matrix - Very large matrices using Python and NumPy

Question

Welcome To Ask or Share your Answers For Others

matrix - Very large matrices using Python and NumPy

1 Reply

深蓝 · Answer 1 · 2021-10-16T22:26:34+0000

PyTables and NumPy are the way to go.

PyTables will store the data on disk in HDF format, with optional compression. My datasets often get 10x compression, which is handy when dealing with tens or hundreds of millions of rows. It's also very fast; my 5 year old laptop can crunch through data doing SQL-like GROUP BY aggregation at 1,000,000 rows/second. Not bad for a Python-based solution!

Accessing the data as a NumPy recarray again is as simple as:

data = table[row_from:row_to]

The HDF library takes care of reading in the relevant chunks of data and converting to NumPy.

Categories

matrix - Very large matrices using Python and NumPy

matrix - Very large matrices using Python and NumPy

Please log in or register to add a comment.

Please log in or register to reply this article.

1 Reply

Please log in or register to add a comment.

Just Browsing Browsing

Most popular tags