Dev Overflow Logo

Dev Overflow

Global search

Search across questions, answers, users and tags.

Loading...
save

Python: fastest way to stream a large CSV into MongoDB?

clock icon

asked 3 months ago

message icon

1

eye icon

418

I have a 6GB CSV to load. Reading it with pandas exhausts memory, and inserting row by row takes hours. What does a reasonable ingest look like?

1 Answer

Stream the file and write in batches — never materialise the whole thing.

1import csv
2from itertools import islice
3from pymongo import MongoClient
4
5BATCH = 5_000
6
7col = MongoClient(uri).devflow.events
8
9with open("big.csv", newline="") as f:
10 reader = csv.DictReader(f)
11 while batch := list(islice(reader, BATCH)):
12 col.insert_many(batch, ordered=False)
1import csv
2from itertools import islice
3from pymongo import MongoClient
4
5BATCH = 5_000
6
7col = MongoClient(uri).devflow.events
8
9with open("big.csv", newline="") as f:
10 reader = csv.DictReader(f)
11 while batch := list(islice(reader, BATCH)):
12 col.insert_many(batch, ordered=False)

ordered=False lets the server keep going past individual duplicate-key failures rather than aborting the batch. Drop secondary indexes before the load and rebuild them afterwards — maintaining them per-insert is usually the real bottleneck.

1

of 1

Write your answer here

Introduce the problem and expand on what you've put in the title.

Top Questions