Editing CommunityData:Hyak Datasets
From CommunityData
Warning: You are not logged in. Your IP address will be publicly visible if you make any edits. If you log in or create an account, your edits will be attributed to your username, along with other benefits.
The edit can be undone. Please check the comparison below to verify that this is what you want to do, and then publish the changes below to finish undoing the edit.
Latest revision | Your text | ||
Line 27: | Line 27: | ||
=== Reading Reddit parquet datasets === | === Reading Reddit parquet datasets === | ||
The recommended way to pull data from parquet on Hyak is to use [https://arrow.apache.org/docs/python/ pyarrow], which makes it relatively easy to filter the data and load it into Pandas. The main alternative is [[CommunityData:Hyak_Spark| Spark]], which is a more complex | The recommended way to pull data from parquet on Hyak is to use [https://arrow.apache.org/docs/python/ pyarrow], which makes it relatively easy to filter the data and load it into Pandas. The main alternative is [[CommunityData:Hyak_Spark| Spark]], which is a more complex system, but can read and write parquet and is useful for working with data that is too large to fit in memory. | ||
This example loads all comments to the Seattle subreddit. You should try it out! | This example loads all comments to the Seattle subreddit. You should try it out! |