Editing CommunityData:Hyak Datasets
From CommunityData
Warning: You are not logged in. Your IP address will be publicly visible if you make any edits. If you log in or create an account, your edits will be attributed to your username, along with other benefits.
The edit can be undone. Please check the comparison below to verify that this is what you want to do, and then publish the changes below to finish undoing the edit.
Latest revision | Your text | ||
Line 28: | Line 28: | ||
The recommended way to pull data from parquet on Hyak is to use [https://arrow.apache.org/docs/python/ pyarrow], which makes it relatively easy to filter the data and load it into Pandas. The main alternative is [[CommunityData:Hyak_Spark| Spark]], which is a more complex and less efficient system, but can read and write parquet and is useful for working with data that is too large to fit in memory. | The recommended way to pull data from parquet on Hyak is to use [https://arrow.apache.org/docs/python/ pyarrow], which makes it relatively easy to filter the data and load it into Pandas. The main alternative is [[CommunityData:Hyak_Spark| Spark]], which is a more complex and less efficient system, but can read and write parquet and is useful for working with data that is too large to fit in memory. | ||
Arrow bindings for R are available, but as of Arrow 0.17.1 it's complicated to install them on Hyak. [[#Install Arrow for R]] | |||
This example loads all comments to the Seattle subreddit. You should try it out! | This example loads all comments to the Seattle subreddit. You should try it out! |