Editing Statistics and Statistical Programming (Winter 2017)/Problem Set: Week 4
From CommunityData
The edit can be undone. Please check the comparison below to verify that this is what you want to do, and then publish the changes below to finish undoing the edit.
Latest revision | Your text | ||
Line 10: | Line 10: | ||
:'''PC2.''' Load both datasets into R as separate data frames. Explore the data to get a sense of the structure of the data. What are the columns, rows, missing data, etc? Write code to take (and then check/look at) several random subsamples of the data. | :'''PC2.''' Load both datasets into R as separate data frames. Explore the data to get a sense of the structure of the data. What are the columns, rows, missing data, etc? Write code to take (and then check/look at) several random subsamples of the data. | ||
:'''PC3.''' Using the top 5000 dataset, create a new data frame where one column is each month (as described in the data) and a second column is the total number of views made to all pages in the dataset over that month. | :'''PC3.''' Using the top 5000 dataset, create a new data frame where one column is each month (as described in the data) and a second column is the total number of views made to all pages in the dataset over that month. | ||
:'''PC4.''' Using the mobile dataset, create a new data frame where one column is each month described in the data and the second is a measure (estimate?) of the total number of views made by mobiles (all platforms) over each month. This will involve at least two steps since total views are included. You'll need to first use the data there to create a measure of the total views | :'''PC4.''' Using the mobile dataset, create a new data frame where one column is each month described in the data and the second is a measure (estimate?) of the total number of views made by mobiles (all platforms) over each month. This will involve at least two steps since total views are included. You'll need to first use the data there to create a measure of the total views per platform. | ||
:'''PC5.''' Merge your two datasets together into a new dataset with columns for each month, total views (across the top 5000 pages) and total mobile views. Are there are missing data? Can you tell why? | :'''PC5.''' Merge your two datasets together into a new dataset with columns for each month, total views (across the top 5000 pages) and total mobile views. Are there are missing data? Can you tell why? | ||
:'''PC6.''' Create a new column in your merged dataset that describes your best estimate of the proportion (or percentage, if you really must!) of views that comes from mobile. Be able to talk about the assumptions you've made here. Make sure that date, in this final column, is a date or datetime object in R. | :'''PC6.''' Create a new column in your merged dataset that describes your best estimate of the proportion (or percentage, if you really must!) of views that comes from mobile. Be able to talk about the assumptions you've made here. Make sure that date, in this final column, is a date or datetime object in R. |