API reference

This page lists everything pycdfpp offers. For explanations and examples, see Reading files and Writing files.

Loading and saving

pycdfpp.load(file_or_buffer: str | PathLike | bytes | bytearray | memoryview, iso_8859_1_to_utf8: bool = True, lazy_load: bool = True)[source]

Load and parse a CDF file.

Parameters:
file_or_bufferstr or os.PathLike or ByteString

Either a file path or an in-memory file implementing the Python buffer protocol.

iso_8859_1_to_utf8bool, optional

Automatically convert Latin-1 characters to their equivalent UTF counterparts when True. For CDF files prior to version 3.8, UTF-8 wasn’t supported and some CDF files might contain “illegal” Latin-1 characters. This option has no impact on valid UTF-8 characters. (Default is True)

lazy_loadbool, optional

Controls whether variable values are loaded immediately or only when accessed by the user. If True, variables’ values are loaded on demand. If False, all variable values are loaded during parsing. (Default is True)

Returns:
CDF
Raises:
FileNotFoundError

When the file doesn’t exist.

ValueError

When the file or buffer is not a valid CDF file.

pycdfpp.save(cdf: CDF, fname: str | PathLike | None = None)[source]

Save a CDF to a file, or to memory.

Saving over the file the CDF was loaded from is safe, even with lazy loading: every value is read before the file is overwritten.

Parameters:
cdfCDF

The CDF to save.

fnamestr or os.PathLike, optional

Destination file name. When omitted, the CDF is serialized in memory.

Returns:
bool or buffer

True when saving to a file; otherwise an object implementing the buffer protocol (e.g. bytes(pycdfpp.save(cdf))).

Raises:
OSError

When the file can’t be written.

Warns:
ExperimentalCompressionWarning

When the CDF or one of its variables uses zstd_compression or blosc2_compression.

Files, variables, attributes

class pycdfpp.CDF

A CDF file object.

Attributes:
attributes: dict

file attributes

variables: dict

file variables

majority: cdf_majority

file majority

distribution_version: int

file distribution version

lazy_loaded: bool

file lazy loading state

compression: CompressionType

file compression type

Methods

add_attribute([name, entries_values, ...])

Adds a new attribute to the CDF.

add_variable([name, values, data_type, ...])

Adds a new variable to the CDF.

add_variable(name=None, values=None, data_type=None, is_nrv=False, compression=CompressionType.no_compression, attributes=None, copy=True) → Variable

Adds a new variable to the CDF.

This method can be called in two ways: 1. With variable parameters: add_variable(name, values=None, data_type=None, is_nrv=False, compression=CompressionType.no_compression, attributes=None, copy=True) 2. With a Variable object: add_variable(variable)

Parameters:
namestr

The name of the variable to add.

valuesnumpy.ndarray or list or None, optional

The values to set for the variable. If None, the variable is created with no values. (Default is None) When a list is passed, the values are converted to a numpy.ndarray with the appropriate data type, with integers, it will choose the smallest data type that can hold all the values.

data_typeDataType or None, optional

The data type of the variable. If None, the data type is inferred from the values. (Default is None)

is_nrvbool, optional

Whether or not the variable is a non-record variable. (Default is False)

compressionCompressionType, optional

The compression type to use for the variable. (Default is CompressionType.no_compression)

attributesMapping[str, List[Any]] or None, optional

The attributes to set for the variable. If None, the variable is created with no attributes. (Default is None)

copybool, optional

If False, the variable borrows the numpy array instead of copying it, see Variable.set_values. Saves the copy of big arrays. (Default is True)

variableVariable

An existing Variable object to add to the CDF (for the second calling method).

Returns:
Variable or None

Returns the newly created variable if successful. Otherwise, returns None.

Raises:
ValueError

If the variable already exists.

Examples

>>> from pycdfpp import CDF, DataType, CompressionType
>>> import numpy as np
>>> cdf = CDF()
>>> # First method: creating a new variable with parameters
>>> cdf.add_variable("var1", np.arange(10, dtype=np.int32), DataType.CDF_INT4, compression=CompressionType.gzip_compression)
var1:
  shape: [ 10 ]
  type: CDF_INT1
  record varry: True
  compression: GNU GZIP
  ...
>>> # Second method: adding an existing variable
>>> cdf2 = CDF()
>>> cdf2.add_variable(cdf["var1"])  # Assuming var1 is already defined in cdf (from the first method)
var1:
  shape: [ 5 ]
  type: CDF_INT1
  record varry: True
  compression: GNU GZIP
  ...
add_attribute(name=None, entries_values=None, entries_types=None) → Attribute

Adds a new attribute to the CDF.

This method can be called in two ways: 1. With attribute parameters: add_attribute(name, entries_values, entries_types=None) 2. With an Attribute object: add_attribute(attribute)

Parameters:
namestr

The name of the attribute to add.

entries_valuesList[np.ndarray or List[float or int or datetime] or str]

The values entries to set for the attribute. When a list is passed, the values are converted to a numpy.ndarray with the appropriate data type, with integers, it will choose the smallest data type that can hold all the values.

entries_typesList[DataType] or None, optional

The data type for each entry of the attribute. If None, the data type is inferred from the values. (Default is None)

attributeAttribute

An existing Attribute object to add to the CDF (for the second calling method).

Returns:
Attribute or None

Returns the newly created attribute if successful. Otherwise, returns None.

Raises:
ValueError

If the attribute already exists.

Examples

>>> from pycdfpp import CDF, DataType
>>> import numpy as np
>>> from datetime import datetime
>>> cdf = CDF()
>>> # First method: creating a new attribute with parameters
>>> cdf.add_attribute("attr1", [np.arange(10, dtype=np.int32)], [DataType.CDF_INT4])
attr1: [ [ [ 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 ] ] ]
>>> # Second method: adding an existing attribute
>>> cdf2 = CDF()
>>> cdf2.add_attribute(cdf.attributes["attr1"])
attr1: [ [ [ 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 ] ] ]
>>> # Another example with multiple entries of different types
>>> cdf.add_attribute("multi", [np.arange(2, dtype=np.int32), [1.,2.,3.], "hello", [datetime(2010,1,1), datetime(2020,1,1)]])
multi: [ [ [ 0, 1 ], [ 1, 2, 3 ], "hello", [ 2010-01-01T00:00:00.000000000, 2020-01-01T00:00:00.000000000 ] ] ]
filter(variables: List[str] | str | Pattern | Callable[[Variable], bool] = None, attributes: List[str] | str | Pattern | Callable[[Attribute], bool] = None, inplace=False) → CDF

Filters the CDF object based on the provided criteria.

Parameters:
cdfCDF

The CDF object to filter.

variablesUnion[List[str], str, re.Pattern, Callable[[Variable], bool]], optional

A list of variable names to keep, a regex pattern, or a callable that returns True for variables to keep. If None (default), all variables are kept.

attributesUnion[List[str], str, re.Pattern, Callable[[Attribute], bool]], optional

A list of global attribute names to keep, a regex pattern, or a callable that returns True for attributes to keep. If None (default), all global attributes are kept.

inplacebool, optional

If True, modifies the original CDF object. If False, returns a new filtered CDF object. (Default is False)

Returns:
CDF

Returns a new CDF object with the filtered variables and attributes.

items(self: pycdfpp._pycdfpp.CDF) → collections.abc.Iterator[tuple[str, pycdfpp._pycdfpp.Variable]]
keys(self: pycdfpp._pycdfpp.CDF) → list[str]
class pycdfpp.Variable

A CDF Variable (either R or Z variable)

Attributes:
attributes: dict

variable attributes

name: str

variable name

type: DataType

variable data type (ie CDF_DOUBLE, CDF_TIME_TT2000, …)

shape: List[int]

variable shape (records + record shape)

majority: cdf_majority

variable majority as writen in the CDF file, note that pycdfpp will always expose row major data.

values_loaded: bool

True if values are availbale in memory, this is usefull with lazy loading to know if values are already loaded.

compression: CompressionType

variable compression type (supported values are no_compression, rle_compression, gzip_compression)

values: numpy.array

returns variable values as a numpy.array of the corresponding dtype and shape, note that no copies are involved, the returned array is just a view on variable data.

values_encoded: numpy.array

same as values except that string variable are encoded wihch involves a data copy and since numpy uses UTF-32, expect a 4x memory increase for string values

Methods

add_attribute([name, values, data_type])

Adds a new attribute to the variable.

set_values(values[, data_type, force, copy])

Sets or resets the values of the variable.

set_compression_type

Sets the variable compression type

add_attribute(name=None, values=None, data_type=None) → VariableAttribute

Adds a new attribute to the variable.

This method can be called in two ways: 1. With attribute parameters: add_attribute(name, values, data_type=None) 2. With a VariableAttribute object: add_attribute(attribute)

Parameters:
namestr

The name of the attribute to add.

valuesnp.ndarray or List[float or int or datetime] or str

The values to set for the attribute. When a list is passed, the values are converted to a numpy.ndarray with the appropriate data type, with integers, it will choose the smallest data type that can hold all the values.

data_typeDataType or None, optional

The data type of the attribute. If None, the data type is inferred from the values. (Default is None)

attributeVariableAttribute

An existing VariableAttribute object to add to the variable (for the second calling method).

Returns:
VariableAttribute

Returns the newly created attribute if successful.

Raises:
ValueError

If the attribute already exists.

Examples

>>> from pycdfpp import CDF, DataType
>>> import numpy as np
>>> cdf = CDF()
>>> cdf.add_variable("var1", np.arange(10, dtype=np.int32), DataType.CDF_INT4)
var1:
  shape: [ 10 ]
  type: CDF_INT1
  record varry: True
  compression: None
  ...
>>> # First method: creating a new attribute with parameters
>>> cdf["var1"].add_attribute("attr1", np.arange(10, dtype=np.int32), DataType.CDF_INT4)
attr1: [ 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 ]
>>> # Second method: adding an existing attribute
>>> var2 = cdf.add_variable("var2", np.arange(5))
>>> var2.add_attribute(cdf["var1"].attributes["attr1"])
attr1: [ 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 ]
set_values(values, data_type=None, force=False, copy=True)

Sets or resets the values of the variable.

Parameters:
valuesnumpy.ndarray or list or tuple or Variable

The values to set for the variable.

data_typeDataType or None, optional

The data type of the variable. If None, the data type is inferred from the values. (Default is None) When passing integer values as a list or tuple, it will choose the smallest data type that can hold all the values. When passing a Variable, the data type is taken from the Variable.

forcebool, optional

If True, allows to overwrite existing values even if the shape or data type do not match. (Default is False)

copybool, optional

If False, the variable borrows the numpy array instead of copying it, and saving writes straight from it: modify the array only if you want the change saved. Reading or modifying the variable’s values copies them first. Raises ValueError if the array can’t be borrowed: it must be C-contiguous, in native byte order, and hold numeric values stored as-is (not strings or times). (Default is True)

Returns:
None
Raises:
ValueError

If the shape or data type do not match and force is False, or if copy is False and the values can’t be borrowed.

Examples

>>> from pycdfpp import CDF, DataType
>>> import numpy as np
>>> cdf = CDF()
>>> cdf.add_variable("var1")
var1:
  shape: [  ]
  type: CDF_NONE
  record vary: True
  compression: None
  ...
>>> # Setting values with numpy array
>>> cdf["var1"].set_values(np.arange(10, 20, dtype=np.int32))
>>> cdf["var1"].values
array([10, 11, 12, 13, 14, 15, 16, 17, 18, 19], dtype=int32)
is_contiguous(self: pycdfpp._pycdfpp.Variable) → bool

Whether the variable’s records are stored as a single contiguous block in the file (True) or fragmented across several VVR/CVVR blocks (False). Walks the variable’s index records on first call, then caches the result.

class pycdfpp.Attribute
type(self: pycdfpp._pycdfpp.Attribute, arg0: SupportsInt) → pycdfpp._pycdfpp.DataType
set_values(entries_values=None, entries_types=None)

Sets the values of the attribute.

This method can be called in two ways: 1. With values and optional types: set_values(entries_values, entries_types=None) 2. With another Attribute object: set_values(attribute)

Parameters:
entries_valuesList[np.ndarray or List[float or int or datetime] or str]

The values entries to set for the attribute. When a list is passed, the values are converted to a numpy.ndarray with the appropriate data type, with integers, it will choose the smallest data type that can hold all the values.

entries_typesList[DataType] or None, optional

The data type for each entry of the attribute. If None, the data type is inferred from the values. (Default is None)

attributeAttribute

An existing Attribute object to set the values from (for the second calling method).

class pycdfpp.VariableAttribute
type(self: pycdfpp._pycdfpp.VariableAttribute) → pycdfpp._pycdfpp.DataType
set_value(value=None, data_type=None)

Sets the value of the variable attribute.

This method can be called in two ways: 1. With value and optional data type: set_value(value, data_type=None) 2. With another VariableAttribute object: set_value(attribute)

Parameters:
valuenp.ndarray or List[float or int or datetime] or str

The value to set for the attribute. When a list is passed, the values are converted to a numpy.ndarray with the appropriate data type, with integers, it will choose the smallest data type that can hold all the values.

data_typeDataType or None, optional

The data type of the attribute. If None, the data type is inferred from the values. (Default is None)

attributeVariableAttribute

An existing VariableAttribute object to set the value from (for the second calling method).

Examples

>>> from pycdfpp import CDF, DataType
>>> import numpy as np
>>> from datetime import datetime
>>> cdf = CDF()
>>> var = cdf.add_variable("var1", np.arange(10, dtype=np.int32), DataType.CDF_INT4)
>>> # First method: setting value with parameters
>>> var.attributes["attr1"].set_value([1, 2, 3])
>>> # Second method: setting from existing attribute
>>> var.attributes["attr2"].set_value(var.attributes["attr1"])
>>> var.attributes["attr2"]
[ 1, 2, 3 ]

Time

pycdfpp.to_datetime64(values)[source]

Convert any compatible given collection of time values to a numpy.datetime64 array.

Parameters:
values: Variable or epoch or List[epoch] or numpy.ndarray[epoch] or epoch16 or List[epoch16] or numpy.array[epoch16] or tt2000_t or List[tt2000_t] or numpy.array[tt2000_t]

input value(s)

to convert to numpy.datetime64
Returns:
numpy.ndarray[numpy.datetime64]
Raises:
TypeError or IndexError

If the input values are not compatible time types.

Notes

On modern x86_64 systems, it will use the CPU’s vectorized instructions to perform the conversion even faster.

pycdfpp.to_datetime(values)[source]

to_datetime

Parameters:
values: Variable or epoch or List[epoch] or epoch16 or List[epoch16] or tt2000_t or List[tt2000_t] or numpy.array[numpy.datetime64[ns]]

input value(s)

to convert to datetime.datetime
Returns:
List[datetime.datetime]
Raises:
TypeError or IndexError

If the input values are not compatible time types.

pycdfpp.to_time_string(values, format: str)[source]

Format CDF time values as an array of fixed-width ASCII strings.

Parameters:
valuesVariable or numpy.ndarray[tt2000_t] or numpy.ndarray[epoch] or numpy.ndarray[epoch16]

CDF time values to format.

formatstr

strftime-compatible format string (e.g. '%Y-%m-%dT%H:%M:%SZ'). %S includes the fraction of a second, with 9 digits (nanoseconds) for every time type.

Returns:
numpy.ndarray

Array of byte strings (dtype S{N}) with the same shape as input.

pycdfpp.to_tt2000(values)[source]

to_tt2000

Parameters:
values: datetime.datetime or List[datetime.datetime] or numpy.array[numpy.datetime64[ns]]

input value(s)

to convert to CDF tt2000
Returns:
tt2000_t or List[tt2000_t]
pycdfpp.to_epoch(values)[source]

to_epoch

Parameters:
values: datetime.datetime or List[datetime.datetime] or numpy.array[numpy.datetime64[ns]]

input value(s)

to convert to CDF epoch
Returns:
epoch or List[epoch]
pycdfpp.to_epoch16(values)[source]

to_epoch16

Parameters:
values: datetime.datetime or List[datetime.datetime] or numpy.array[numpy.datetime64[ns]]

input value(s)

to convert to CDF epoch16
Returns:
epoch16 or List[epoch16]
class pycdfpp.tt2000_t
class pycdfpp.epoch
class pycdfpp.epoch16

Helpers

pycdfpp.default_fill_value(cdf_type: DataType)[source]

Return a default fill value for the given CDF data type.

Parameters:
cdf_typeDataType

The CDF data type for which to return the default fill value.

Returns
——-
Any

The default fill value for the specified CDF data type.

pycdfpp.default_pad_value(cdf_type: DataType)[source]

Returns the default pad value for the given CDF data type (CDF User’s Guide, table 2.8): the value of records a file doesn’t store, when it declares no pad value of its own.

pycdfpp.to_dict_skeleton(obj: Any) → Any[source]
pycdfpp.to_dict_skeleton(attribute: Attribute) → dict
pycdfpp.to_dict_skeleton(attribute: VariableAttribute) → dict
pycdfpp.to_dict_skeleton(variable: Variable) → dict
pycdfpp.to_dict_skeleton(cdf: CDF) → dict

Enumerations

class pycdfpp.DataType(*values)
CDF_BYTE = 41
CDF_CHAR = 51
CDF_INT1 = 1
CDF_INT2 = 2
CDF_INT4 = 4
CDF_INT8 = 8
CDF_NONE = 0
CDF_EPOCH = 31
CDF_FLOAT = 44
CDF_REAL4 = 21
CDF_REAL8 = 22
CDF_UCHAR = 52
CDF_UINT1 = 11
CDF_UINT2 = 12
CDF_UINT4 = 14
CDF_DOUBLE = 45
CDF_EPOCH16 = 32
CDF_TIME_TT2000 = 33
class pycdfpp.CompressionType(*values)
no_compression = 0
gzip_compression = 5
rle_compression = 1
ahuff_compression = 3
huff_compression = 2
zstd_compression = 16
blosc2_compression = 17
class pycdfpp.Majority(*values)
row = 1
column = 0

Warnings

class pycdfpp.ExperimentalCompressionWarning[source]

A CDF was saved with a codec outside the CDF standard (zstd, blosc2): only CDFpp can read it.

Low-level inspection

pycdfpp.debug

Structured access to CDFpp’s on-disk record structure - the physical layer underneath load()’s reconstructed Variable/Attribute view. Useful for building tools like cdfdump, or diagnosing a file load() itself can’t fully parse.

pycdfpp.debug.for_each_record()

debug_for_each_record(path: str) -> list

Walk a CDF file’s records in physical disk order (record_size stepping from one header to the next), not the reconstructed logical variable/attribute graph load() gives you. Surfaces records the semantic loader silently discards (UIR - freed space CDF leaves behind rather than compacting) and tolerates record types not yet modeled (SPR) by skipping them via their declared size.

Parameters:
pathstr

Path to the CDF file.

Returns:
list of tuple[int, str, dict]

One entry per on-disk record, in file order: (byte offset, record type name, {field_name: value}). A structurally corrupted record aborts the walk after printing a diagnostic to stderr (same default policy as the C++ API).

pycdfpp.debug.nasa_compat_dump(path: str, radix: SupportsInt = 10, show_data: bool = False, summary: bool = False) → str

Dump a CDF file’s records in NASA’s own cdfirsdump (-full -nopage) text format, byte-for-byte (verified against real captures of that tool’s own output - see tests/nasa_compat_repr). Built on the same physical-order walk as for_each_record, not load()’s reconstructed view.

Parameters:
pathstr

Path to the CDF file.

radixint, default 10

10 (decimal) or 16 (hex, “0x” + 16 uppercase digits) for record offsets.

show_databool, default False

Hex-dump VVR/CVVR payload bytes (cdfirsdump’s -data).

summarybool, default False

Append the closing record-type summary table (cdfirsdump’s default -summary behavior; this binding defaults to False to match this project’s own existing default, not NASA’s).

Returns:
str

The full dump text, ready to print or write to a file.

pycdfpp.debug.nasa_compat_dump_from_offset(path: str, offset: SupportsInt, radix: SupportsInt = 10, show_data: bool = False) → str

Same as nasa_compat_dump, but starts the walk at a given byte offset instead of the file start - no magic-number preamble is printed (matching a real cdfirsdump -offset capture).

Parameters:
pathstr

Path to the CDF file.

offsetint

Byte offset to start the walk at.

radixint, default 10

10 (decimal) or 16 (hex) for record offsets.

show_databool, default False

Hex-dump VVR/CVVR payload bytes.

Returns:
str

The dump text from that offset onward.

pycdfpp.debug.nasa_compat_dump_brief(path: str) → str

Dump only the record-type summary table (cdfirsdump’s -brief, the real tool’s own default level), byte-for-byte, including its “(+16 if with checksum)” quirk which is always shown at brief level regardless of the file (see nasa_compat_repr.hpp’s own notes on why).

Parameters:
pathstr

Path to the CDF file.

Returns:
str

The banner + summary table text.