Editing Human Centered Data Science (Fall 2019)/Assignments
From CommunityData
Warning: You are not logged in. Your IP address will be publicly visible if you make any edits. If you log in or create an account, your edits will be attributed to your username, along with other benefits.
The edit can be undone. Please check the comparison below to verify that this is what you want to do, and then publish the changes below to finish undoing the edit.
Latest revision | Your text | ||
Line 278: | Line 278: | ||
Some processing of the data will be necessary! In particular, you'll need to - after retrieving and including the ORES data for each article - merge the wikipedia data and population data together. Both have fields containing country names for just that purpose. After merging the data, you'll invariably run into entries which ''cannot'' be merged. Either the population dataset does not have an entry for the equivalent Wikipedia country, or vis versa. | Some processing of the data will be necessary! In particular, you'll need to - after retrieving and including the ORES data for each article - merge the wikipedia data and population data together. Both have fields containing country names for just that purpose. After merging the data, you'll invariably run into entries which ''cannot'' be merged. Either the population dataset does not have an entry for the equivalent Wikipedia country, or vis versa. | ||
Please remove any rows that do not have matching data, and output them to a CSV file called <tt> | Please remove any rows that do not have matching data, and output them to a CSV file called <tt>population_by_country-no_match.csv</tt> | ||
Consolidate the remaining data into a single CSV file | Consolidate the remaining data into a single CSV file which looks something like this: | ||