Performance¶
How does pycdfpp compare with the two other Python CDF readers,
spacepy.pycdf and
cdflib? This page measures everyday tasks on real
files, then explains where the differences come from.
The short answer¶
pycdfpp is faster on every task we measured but one. When you read whole files of
compressed data, it is 5.6× to 15× faster. When you open a file, or convert time
variables, it is 11× to about 3700× faster. Writing compressed files is 3.7× to 14×
faster. Writing without compression, spacepy is 13% faster.
Results¶
Task |
Data |
pycdfpp |
spacepy.pycdf |
cdflib |
|---|---|---|---|---|
Open a file, list variables, read all attributes |
MMS FPI electron distribution, 178 MB |
0.6 ms |
264 ms (407×) |
9.9 ms (15×) |
Read B and its time axis as |
MMS FGM survey, 1.2 M points, gzip, TT2000 |
7.8 ms |
3.77 s (483×) |
306 ms (39×) |
Read B and its time axis as |
Wind MFI, 0.9 M points, CDF_EPOCH |
3.6 ms |
40.2 ms (11×) |
13.3 s (3718×) |
Read every variable of a file |
MMS FPI electron distribution, 178 MB, gzip |
110 ms |
960 ms (8.7×) |
620 ms (5.6×) |
Read every variable of a folder |
23 CDAWeb files, 11 missions, 528 MB |
391 ms |
2.37 s (6.1×) |
2.81 s (7.2×) |
Same folder, 8 threads |
23 CDAWeb files, 11 missions, 528 MB |
161 ms |
not thread-safe |
2.41 s (15×) |
Write B and its time axis, gzip |
MMS FGM survey, 1.2 M points, 29 MB |
46.0 ms |
458 ms (10×) |
172 ms (3.7×) |
Write B and its time axis, uncompressed |
MMS FGM survey, 1.2 M points, 29 MB |
14.1 ms |
12.3 ms (0.9×) |
17.4 ms (1.2×) |
Write a particle distribution file, gzip |
MMS FPI electron distribution, 210 MB |
275 ms |
3.93 s (14×) |
1.87 s (6.8×) |
(N×) means N times longer than pycdfpp. Each time is the median of 5 runs, after
one warm-up run, so the files are in the page cache.
Measured on an AMD Ryzen 7 5800X (8 cores, AVX2, no AVX-512), Linux, Python 3.13, with the packages from PyPI: pycdfpp 0.15.0, spacepy 0.7.0 (which bundles NASA’s CDF library 3.9.0), cdflib 1.3.14 and numpy 2.5.3.
How it was measured¶
The script is benchmarks/python_libs/compare.py. It downloads the files from CDAWeb once, then caches them. Run it yourself:
$ pip install pycdfpp spacepy cdflib requests
$ python benchmarks/python_libs/compare.py
Each library is used the way its documentation shows, with its fastest documented
way to get datetime64:
pycdfpp:pycdfpp.load(path),cdf[name].valuesandpycdfpp.to_datetime64(cdf[name]).spacepy:pycdf.CDF(path)andcdf.raw_var(name)[...]for values. For CDF_EPOCH times,spacepy.time.Ticktock(raw, "CDF").UNX, which is vectorized. Ticktock doesn’t handle TT2000, so TT2000 times go throughcdf[name][...].cdflib:cdflib.CDF(path),cdf.varget(name)andcdflib.cdfepoch.to_datetime(...).
Writing, each library writes the same variables, with gzip level 6 when compressed:
pycdfpp.save of a CDF built with add_variable, borrowing the data arrays
with copy=False; spacepy’s cdf.new and
raw_var with TT2000 integers; cdflib’s CDFWriter. Every written file is read
back and checked.
We checked that the three libraries return the same values for every variable of the 23 files.
Why is pycdfpp faster?¶
Each row of the table has its own reason.
Opening a file¶
pycdfppmaps the file in memory and parses only its headers, in C++. Variable data is read later, when you ask for it.NASA’s library, used by
spacepy, checks the file’s MD5 checksum every time it opens a file that has one. MMS files all have one.Checking the checksum means reading and hashing the whole file. For 178 MB, that is about 250 ms, before you have read anything.
cdflibparses the headers in Python, which takes about 10 ms.
pycdfpp does not check checksums. If you need that check, use NASA’s tools.
Converting time¶
pycdfppconverts CDF time values todatetime64[ns]in C++, using SIMD instructions, and exactly. On the test machine (AVX2), it converts about one billion TT2000 values and 2.6 billion CDF_EPOCH values per second.For TT2000,
spacepycreates one Pythondatetimeobject per value. That takes seconds for a million points.datetimealso stops at microseconds, so nanoseconds are lost. For CDF_EPOCH,spacepy.time.Ticktockis vectorized, so the gap is much smaller.cdflibconverts TT2000 with numpy, which is reasonably fast. But it converts CDF_EPOCH in a Python loop, one value at a time. That is why the Wind file takes 14 s.
CDF_EPOCH is a plain count of milliseconds, with no leap seconds. So you can also convert it yourself with one line of numpy, whatever the library. TT2000 needs a leap-second table, which is why a library function matters more there.
pycdfpp also keeps the fractions of milliseconds that CDF_EPOCH values can hold.
cdflib rounds them down to the millisecond.
Decompressing¶
Most mission files compress their variables with gzip, in many blocks: an MMS FPI
distribution variable has 640 of them. Two things make pycdfpp faster there:
It decompresses the blocks of a variable on all cores at once. The two other libraries decompress them one after another.
It uses libdeflate, the two others zlib. On these files, libdeflate alone decompresses 1.5 to 1.7 times faster.
On Linux, big buffers use 2 MB huge pages: filling them on one thread is 2 to 3 times faster. Before the threads start, one thread touches every page. Otherwise several threads fault the same fresh huge page at once, and the kernel zeroes one for each of them.
Decompressing takes most of the time for big compressed files, so these rows gain the most from it: 5.6× to 8.7× here, up to 15× with threads.
Using threads¶
pycdfppreleases Python’s GIL while it reads and decompresses. So several threads really run at the same time.With 8 threads, the folder loads in 0.16 s instead of 0.39 s. The biggest file alone takes 0.11 s, even with its own blocks decompressed in parallel.
cdflibruns mostly Python code, which holds the GIL. Threads barely help it.NASA’s library keeps global state, so
spacepycan’t be used from several threads.
See Reading files for how to load many files with a thread pool.
Writing¶
pycdfppcuts compressed variables into 256 KB blocks and compresses them on all cores, with libdeflate.NASA’s library, used by
spacepy, andcdflibboth compress with zlib, on one thread.Without compression, writing is mostly copying memory to the file. With
copy=False(see Writing files),pycdfppwrites straight from the arrays, likespacepy. The rest of the gap is the time axis:pycdfppconverts it fromdatetime64, which takes 1.3 ms here, whilespacepyis given TT2000 integers.
Scaling¶
The script benchmarks/python_libs/scaling.py measures how each library scales with threads and with file size. Same machine and versions as above.
Reading the 23-file folder with a thread pool (spacepy can’t use threads):
Threads |
pycdfpp |
cdflib |
|---|---|---|
1 |
0.35 s |
2.79 s |
2 |
0.21 s (1.6×) |
2.31 s (1.2×) |
4 |
0.16 s (2.2×) |
2.27 s (1.2×) |
8 |
0.16 s (2.3×) |
2.53 s (1.1×) |
16 |
0.16 s (2.2×) |
2.47 s (1.1×) |
(N×) is the speed-up over one thread. pycdfpp stops at about 2.3× because one file
of the folder takes 0.11 s on its own.
Writing and reading one float32 variable of 3 components, gzip compressed, in MB/s
of values (higher is better). Every library reads the file written by NASA’s library:
Size |
Write: pycdfpp |
spacepy |
cdflib |
Read: pycdfpp |
spacepy |
cdflib |
|---|---|---|---|---|---|---|
1 MB |
111 |
41 |
112 |
498 |
277 |
453 |
10 MB |
606 |
42 |
113 |
4085 |
273 |
438 |
100 MB |
697 |
42 |
104 |
3450 |
270 |
317 |
1000 MB |
712 |
42 |
103 |
3550 |
264 |
326 |
At 1 MB, a variable is too small to be worth several threads, so
pycdfppcompresses and decompresses on one thread, like the others.From 10 MB up, it compresses and decompresses on all cores, and keeps that speed up to 1 GB.