pandas - Python: Unstacked DataFrame is too big, causing int32 overflow

Question

Ask a Question

Welcome To Ask or Share your Answers For Others

pandas - Python: Unstacked DataFrame is too big, causing int32 overflow

asked Oct 24, 2021 in Technique[技术] by 深蓝 (71.8m points)

I have a big dataset and when I try to run this code I get a memory error.

user_by_movie = user_items.groupby(['user_id', 'movie_id'])['rating'].max().unstack()

here is the error:

ValueError: Unstacked DataFrame is too big, causing int32 overflow

I have run it on another machine and it worked fine! how can I fix this error?

See Question&Answers more detail:os

与恶龙缠斗过久,自身亦成为恶龙；凝视深渊过久,深渊将回以凝视…

3.9k views

1 Answer

深蓝 · Answer 1 · 2021-10-23T21:35:33+0000

As it turns out this was not an issue on pandas 0.21. I am using a Jupyter notebook and I need the latest version of pandas for the rest of the code. So I did this:

!pip install pandas==0.21
import pandas as pd
user_by_movie = user_items.groupby(['user_id', 'movie_id'])['rating'].max().unstack()
!pip install pandas

This code works on the Jupyter notebook. First, it downgrades pandas to 0.21 and runs the code. After having the required dataset it updates pandas to the latest version. check the issue raised on GitHub here. This post was also helpful to increase memory of Jupyter notebook.

Categories

pandas - Python: Unstacked DataFrame is too big, causing int32 overflow

Please log in or register to add a comment.

Please log in or register to answer this question.

1 Answer

Please log in or register to add a comment.

Just Browsing Browsing

Most popular tags