Editing Human Centered Data Science (Fall 2019)/Assignments
From CommunityData
Warning: You are not logged in. Your IP address will be publicly visible if you make any edits. If you log in or create an account, your edits will be attributed to your username, along with other benefits.
The edit can be undone. Please check the comparison below to verify that this is what you want to do, and then publish the changes below to finish undoing the edit.
Latest revision | Your text | ||
Line 276: | Line 276: | ||
==== Combining the datasets ==== | ==== Combining the datasets ==== | ||
Some processing of the data will be necessary! In particular, you'll need to - after retrieving and including the ORES data for each article - merge the wikipedia data and population data together. Both have fields containing country names for just that purpose. After merging the data, you'll invariably run into entries which ''cannot'' be merged. Either the population dataset does not have an entry for the equivalent Wikipedia country, or | Some processing of the data will be necessary! In particular, you'll need to - after retrieving and including the ORES data for each article - merge the wikipedia data and population data together. Both have fields containing country names for just that purpose. After merging the data, you'll invariably run into entries which ''cannot'' be merged. Either the population dataset does not have an entry for the equivalent Wikipedia country, or vice versa. | ||
Please remove any rows that do not have matching data, and output them to a CSV file called <tt>wp_wpds_countries-no_match.csv</tt> | Please remove any rows that do not have matching data, and output them to a CSV file called <tt>wp_wpds_countries-no_match.csv</tt> |