API reference¶
This page lists everything pycdfpp offers. For explanations and examples, see
Reading files and Writing files.
Loading and saving¶
- pycdfpp.load(file_or_buffer: str | PathLike | bytes | bytearray | memoryview, iso_8859_1_to_utf8: bool = True, lazy_load: bool = True)[source]¶
Load and parse a CDF file.
- Parameters:
- file_or_bufferstr or os.PathLike or ByteString
Either a file path or an in-memory file implementing the Python buffer protocol.
- iso_8859_1_to_utf8bool, optional
Automatically convert Latin-1 characters to their equivalent UTF counterparts when True. For CDF files prior to version 3.8, UTF-8 wasn’t supported and some CDF files might contain “illegal” Latin-1 characters. This option has no impact on valid UTF-8 characters. (Default is True)
- lazy_loadbool, optional
Controls whether variable values are loaded immediately or only when accessed by the user. If True, variables’ values are loaded on demand. If False, all variable values are loaded during parsing. (Default is True)
- Returns:
- CDF
- Raises:
- FileNotFoundError
When the file doesn’t exist.
- ValueError
When the file or buffer is not a valid CDF file.
- pycdfpp.save(cdf: CDF, fname: str | PathLike | None = None)[source]¶
Save a CDF to a file, or to memory.
Saving over the file the CDF was loaded from is safe, even with lazy loading: every value is read before the file is overwritten.
- Parameters:
- cdfCDF
The CDF to save.
- fnamestr or os.PathLike, optional
Destination file name. When omitted, the CDF is serialized in memory.
- Returns:
- bool or buffer
True when saving to a file; otherwise an object implementing the buffer protocol (e.g.
bytes(pycdfpp.save(cdf))).
- Raises:
- OSError
When the file can’t be written.
- Warns:
- ExperimentalCompressionWarning
When the CDF or one of its variables uses zstd_compression or blosc2_compression.
Files, variables, attributes¶
- class pycdfpp.CDF¶
A CDF file object.
- Attributes:
- attributes: dict
file attributes
- variables: dict
file variables
- majority: cdf_majority
file majority
- distribution_version: int
file distribution version
- lazy_loaded: bool
file lazy loading state
- compression: CompressionType
file compression type
Methods
add_attribute([name, entries_values, ...])Adds a new attribute to the CDF.
add_variable([name, values, data_type, ...])Adds a new variable to the CDF.
- add_variable(name=None, values=None, data_type=None, is_nrv=False, compression=CompressionType.no_compression, attributes=None, copy=True) Variable¶
Adds a new variable to the CDF.
This method can be called in two ways: 1. With variable parameters: add_variable(name, values=None, data_type=None, is_nrv=False, compression=CompressionType.no_compression, attributes=None, copy=True) 2. With a Variable object: add_variable(variable)
- Parameters:
- namestr
The name of the variable to add.
- valuesnumpy.ndarray or list or None, optional
The values to set for the variable. If None, the variable is created with no values. (Default is None) When a list is passed, the values are converted to a numpy.ndarray with the appropriate data type, with integers, it will choose the smallest data type that can hold all the values.
- data_typeDataType or None, optional
The data type of the variable. If None, the data type is inferred from the values. (Default is None)
- is_nrvbool, optional
Whether or not the variable is a non-record variable. (Default is False)
- compressionCompressionType, optional
The compression type to use for the variable. (Default is CompressionType.no_compression)
- attributesMapping[str, List[Any]] or None, optional
The attributes to set for the variable. If None, the variable is created with no attributes. (Default is None)
- copybool, optional
If False, the variable borrows the numpy array instead of copying it, see Variable.set_values. Saves the copy of big arrays. (Default is True)
- variableVariable
An existing Variable object to add to the CDF (for the second calling method).
- Returns:
- Variable or None
Returns the newly created variable if successful. Otherwise, returns None.
- Raises:
- ValueError
If the variable already exists.
Examples
>>> from pycdfpp import CDF, DataType, CompressionType >>> import numpy as np >>> cdf = CDF() >>> # First method: creating a new variable with parameters >>> cdf.add_variable("var1", np.arange(10, dtype=np.int32), DataType.CDF_INT4, compression=CompressionType.gzip_compression) var1: shape: [ 10 ] type: CDF_INT1 record varry: True compression: GNU GZIP ... >>> # Second method: adding an existing variable >>> cdf2 = CDF() >>> cdf2.add_variable(cdf["var1"]) # Assuming var1 is already defined in cdf (from the first method) var1: shape: [ 5 ] type: CDF_INT1 record varry: True compression: GNU GZIP ...
- add_attribute(name=None, entries_values=None, entries_types=None) Attribute¶
Adds a new attribute to the CDF.
This method can be called in two ways: 1. With attribute parameters: add_attribute(name, entries_values, entries_types=None) 2. With an Attribute object: add_attribute(attribute)
- Parameters:
- namestr
The name of the attribute to add.
- entries_valuesList[np.ndarray or List[float or int or datetime] or str]
The values entries to set for the attribute. When a list is passed, the values are converted to a numpy.ndarray with the appropriate data type, with integers, it will choose the smallest data type that can hold all the values.
- entries_typesList[DataType] or None, optional
The data type for each entry of the attribute. If None, the data type is inferred from the values. (Default is None)
- attributeAttribute
An existing Attribute object to add to the CDF (for the second calling method).
- Returns:
- Attribute or None
Returns the newly created attribute if successful. Otherwise, returns None.
- Raises:
- ValueError
If the attribute already exists.
Examples
>>> from pycdfpp import CDF, DataType >>> import numpy as np >>> from datetime import datetime >>> cdf = CDF() >>> # First method: creating a new attribute with parameters >>> cdf.add_attribute("attr1", [np.arange(10, dtype=np.int32)], [DataType.CDF_INT4]) attr1: [ [ [ 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 ] ] ] >>> # Second method: adding an existing attribute >>> cdf2 = CDF() >>> cdf2.add_attribute(cdf.attributes["attr1"]) attr1: [ [ [ 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 ] ] ] >>> # Another example with multiple entries of different types >>> cdf.add_attribute("multi", [np.arange(2, dtype=np.int32), [1.,2.,3.], "hello", [datetime(2010,1,1), datetime(2020,1,1)]]) multi: [ [ [ 0, 1 ], [ 1, 2, 3 ], "hello", [ 2010-01-01T00:00:00.000000000, 2020-01-01T00:00:00.000000000 ] ] ]
- filter(variables: List[str] | str | Pattern | Callable[[Variable], bool] = None, attributes: List[str] | str | Pattern | Callable[[Attribute], bool] = None, inplace=False) CDF¶
Filters the CDF object based on the provided criteria.
- Parameters:
- cdfCDF
The CDF object to filter.
- variablesUnion[List[str], str, re.Pattern, Callable[[Variable], bool]], optional
A list of variable names to keep, a regex pattern, or a callable that returns True for variables to keep. If None (default), all variables are kept.
- attributesUnion[List[str], str, re.Pattern, Callable[[Attribute], bool]], optional
A list of global attribute names to keep, a regex pattern, or a callable that returns True for attributes to keep. If None (default), all global attributes are kept.
- inplacebool, optional
If True, modifies the original CDF object. If False, returns a new filtered CDF object. (Default is False)
- Returns:
- CDF
Returns a new CDF object with the filtered variables and attributes.
- items(self: pycdfpp._pycdfpp.CDF) collections.abc.Iterator[tuple[str, pycdfpp._pycdfpp.Variable]]¶
- keys(self: pycdfpp._pycdfpp.CDF) list[str]¶
- class pycdfpp.Variable¶
A CDF Variable (either R or Z variable)
- Attributes:
- attributes: dict
variable attributes
- name: str
variable name
- type: DataType
variable data type (ie CDF_DOUBLE, CDF_TIME_TT2000, …)
- shape: List[int]
variable shape (records + record shape)
- majority: cdf_majority
variable majority as writen in the CDF file, note that pycdfpp will always expose row major data.
- values_loaded: bool
True if values are availbale in memory, this is usefull with lazy loading to know if values are already loaded.
- compression: CompressionType
variable compression type (supported values are no_compression, rle_compression, gzip_compression)
- values: numpy.array
returns variable values as a numpy.array of the corresponding dtype and shape, note that no copies are involved, the returned array is just a view on variable data.
- values_encoded: numpy.array
same as values except that string variable are encoded wihch involves a data copy and since numpy uses UTF-32, expect a 4x memory increase for string values
Methods
add_attribute([name, values, data_type])Adds a new attribute to the variable.
set_values(values[, data_type, force, copy])Sets or resets the values of the variable.
set_compression_type
Sets the variable compression type
- add_attribute(name=None, values=None, data_type=None) VariableAttribute¶
Adds a new attribute to the variable.
This method can be called in two ways: 1. With attribute parameters: add_attribute(name, values, data_type=None) 2. With a VariableAttribute object: add_attribute(attribute)
- Parameters:
- namestr
The name of the attribute to add.
- valuesnp.ndarray or List[float or int or datetime] or str
The values to set for the attribute. When a list is passed, the values are converted to a numpy.ndarray with the appropriate data type, with integers, it will choose the smallest data type that can hold all the values.
- data_typeDataType or None, optional
The data type of the attribute. If None, the data type is inferred from the values. (Default is None)
- attributeVariableAttribute
An existing VariableAttribute object to add to the variable (for the second calling method).
- Returns:
- VariableAttribute
Returns the newly created attribute if successful.
- Raises:
- ValueError
If the attribute already exists.
Examples
>>> from pycdfpp import CDF, DataType >>> import numpy as np >>> cdf = CDF() >>> cdf.add_variable("var1", np.arange(10, dtype=np.int32), DataType.CDF_INT4) var1: shape: [ 10 ] type: CDF_INT1 record varry: True compression: None ... >>> # First method: creating a new attribute with parameters >>> cdf["var1"].add_attribute("attr1", np.arange(10, dtype=np.int32), DataType.CDF_INT4) attr1: [ 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 ] >>> # Second method: adding an existing attribute >>> var2 = cdf.add_variable("var2", np.arange(5)) >>> var2.add_attribute(cdf["var1"].attributes["attr1"]) attr1: [ 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 ]
- set_values(values, data_type=None, force=False, copy=True)¶
Sets or resets the values of the variable.
- Parameters:
- valuesnumpy.ndarray or list or tuple or Variable
The values to set for the variable.
- data_typeDataType or None, optional
The data type of the variable. If None, the data type is inferred from the values. (Default is None) When passing integer values as a list or tuple, it will choose the smallest data type that can hold all the values. When passing a Variable, the data type is taken from the Variable.
- forcebool, optional
If True, allows to overwrite existing values even if the shape or data type do not match. (Default is False)
- copybool, optional
If False, the variable borrows the numpy array instead of copying it, and saving writes straight from it: modify the array only if you want the change saved. Reading or modifying the variable’s values copies them first. Raises ValueError if the array can’t be borrowed: it must be C-contiguous, in native byte order, and hold numeric values stored as-is (not strings or times). (Default is True)
- Returns:
- None
- Raises:
- ValueError
If the shape or data type do not match and force is False, or if copy is False and the values can’t be borrowed.
Examples
>>> from pycdfpp import CDF, DataType >>> import numpy as np >>> cdf = CDF() >>> cdf.add_variable("var1") var1: shape: [ ] type: CDF_NONE record vary: True compression: None ... >>> # Setting values with numpy array >>> cdf["var1"].set_values(np.arange(10, 20, dtype=np.int32)) >>> cdf["var1"].values array([10, 11, 12, 13, 14, 15, 16, 17, 18, 19], dtype=int32)
- is_contiguous(self: pycdfpp._pycdfpp.Variable) bool¶
Whether the variable’s records are stored as a single contiguous block in the file (True) or fragmented across several VVR/CVVR blocks (False). Walks the variable’s index records on first call, then caches the result.
- class pycdfpp.Attribute¶
- type(self: pycdfpp._pycdfpp.Attribute, arg0: SupportsInt) pycdfpp._pycdfpp.DataType¶
- set_values(entries_values=None, entries_types=None)¶
Sets the values of the attribute.
This method can be called in two ways: 1. With values and optional types: set_values(entries_values, entries_types=None) 2. With another Attribute object: set_values(attribute)
- Parameters:
- entries_valuesList[np.ndarray or List[float or int or datetime] or str]
The values entries to set for the attribute. When a list is passed, the values are converted to a numpy.ndarray with the appropriate data type, with integers, it will choose the smallest data type that can hold all the values.
- entries_typesList[DataType] or None, optional
The data type for each entry of the attribute. If None, the data type is inferred from the values. (Default is None)
- attributeAttribute
An existing Attribute object to set the values from (for the second calling method).
- class pycdfpp.VariableAttribute¶
-
- set_value(value=None, data_type=None)¶
Sets the value of the variable attribute.
This method can be called in two ways: 1. With value and optional data type: set_value(value, data_type=None) 2. With another VariableAttribute object: set_value(attribute)
- Parameters:
- valuenp.ndarray or List[float or int or datetime] or str
The value to set for the attribute. When a list is passed, the values are converted to a numpy.ndarray with the appropriate data type, with integers, it will choose the smallest data type that can hold all the values.
- data_typeDataType or None, optional
The data type of the attribute. If None, the data type is inferred from the values. (Default is None)
- attributeVariableAttribute
An existing VariableAttribute object to set the value from (for the second calling method).
Examples
>>> from pycdfpp import CDF, DataType >>> import numpy as np >>> from datetime import datetime >>> cdf = CDF() >>> var = cdf.add_variable("var1", np.arange(10, dtype=np.int32), DataType.CDF_INT4) >>> # First method: setting value with parameters >>> var.attributes["attr1"].set_value([1, 2, 3]) >>> # Second method: setting from existing attribute >>> var.attributes["attr2"].set_value(var.attributes["attr1"]) >>> var.attributes["attr2"] [ 1, 2, 3 ]
Time¶
- pycdfpp.to_datetime64(values)[source]¶
Convert any compatible given collection of time values to a numpy.datetime64 array.
- Parameters:
- values: Variable or epoch or List[epoch] or numpy.ndarray[epoch] or epoch16 or List[epoch16] or numpy.array[epoch16] or tt2000_t or List[tt2000_t] or numpy.array[tt2000_t]
input value(s)
- to convert to numpy.datetime64
- Returns:
- numpy.ndarray[numpy.datetime64]
- Raises:
- TypeError or IndexError
If the input values are not compatible time types.
Notes
On modern x86_64 systems, it will use the CPU’s vectorized instructions to perform the conversion even faster.
- pycdfpp.to_datetime(values)[source]¶
to_datetime
- Parameters:
- values: Variable or epoch or List[epoch] or epoch16 or List[epoch16] or tt2000_t or List[tt2000_t] or numpy.array[numpy.datetime64[ns]]
input value(s)
- to convert to datetime.datetime
- Returns:
- List[datetime.datetime]
- Raises:
- TypeError or IndexError
If the input values are not compatible time types.
- pycdfpp.to_time_string(values, format: str)[source]¶
Format CDF time values as an array of fixed-width ASCII strings.
- Parameters:
- valuesVariable or numpy.ndarray[tt2000_t] or numpy.ndarray[epoch] or numpy.ndarray[epoch16]
CDF time values to format.
- formatstr
strftime-compatible format string (e.g.
'%Y-%m-%dT%H:%M:%SZ').%Sincludes the fraction of a second, with 9 digits (nanoseconds) for every time type.
- Returns:
- numpy.ndarray
Array of byte strings (dtype
S{N}) with the same shape as input.
- pycdfpp.to_tt2000(values)[source]¶
to_tt2000
- Parameters:
- values: datetime.datetime or List[datetime.datetime] or numpy.array[numpy.datetime64[ns]]
input value(s)
- to convert to CDF tt2000
- Returns:
- tt2000_t or List[tt2000_t]
- pycdfpp.to_epoch(values)[source]¶
to_epoch
- Parameters:
- values: datetime.datetime or List[datetime.datetime] or numpy.array[numpy.datetime64[ns]]
input value(s)
- to convert to CDF epoch
- Returns:
- epoch or List[epoch]
- pycdfpp.to_epoch16(values)[source]¶
to_epoch16
- Parameters:
- values: datetime.datetime or List[datetime.datetime] or numpy.array[numpy.datetime64[ns]]
input value(s)
- to convert to CDF epoch16
- Returns:
- epoch16 or List[epoch16]
- class pycdfpp.tt2000_t¶
- class pycdfpp.epoch¶
- class pycdfpp.epoch16¶
Helpers¶
- pycdfpp.default_fill_value(cdf_type: DataType)[source]¶
Return a default fill value for the given CDF data type.
- Parameters:
- cdf_typeDataType
The CDF data type for which to return the default fill value.
- Returns
- ——-
- Any
The default fill value for the specified CDF data type.
Enumerations¶
- class pycdfpp.DataType(*values)¶
- CDF_BYTE = 41¶
- CDF_CHAR = 51¶
- CDF_INT1 = 1¶
- CDF_INT2 = 2¶
- CDF_INT4 = 4¶
- CDF_INT8 = 8¶
- CDF_NONE = 0¶
- CDF_EPOCH = 31¶
- CDF_FLOAT = 44¶
- CDF_REAL4 = 21¶
- CDF_REAL8 = 22¶
- CDF_UCHAR = 52¶
- CDF_UINT1 = 11¶
- CDF_UINT2 = 12¶
- CDF_UINT4 = 14¶
- CDF_DOUBLE = 45¶
- CDF_EPOCH16 = 32¶
- CDF_TIME_TT2000 = 33¶
Warnings¶
Low-level inspection¶
pycdfpp.debug¶
Structured access to CDFpp’s on-disk record structure - the physical layer
underneath load()’s reconstructed Variable/Attribute view. Useful for building
tools like cdfdump, or diagnosing a file load() itself can’t fully parse.
- pycdfpp.debug.for_each_record()¶
debug_for_each_record(path: str) -> list
Walk a CDF file’s records in physical disk order (record_size stepping from one header to the next), not the reconstructed logical variable/attribute graph load() gives you. Surfaces records the semantic loader silently discards (UIR - freed space CDF leaves behind rather than compacting) and tolerates record types not yet modeled (SPR) by skipping them via their declared size.
- Parameters:
- pathstr
Path to the CDF file.
- Returns:
- list of tuple[int, str, dict]
One entry per on-disk record, in file order: (byte offset, record type name, {field_name: value}). A structurally corrupted record aborts the walk after printing a diagnostic to stderr (same default policy as the C++ API).
- pycdfpp.debug.nasa_compat_dump(path: str, radix: SupportsInt = 10, show_data: bool = False, summary: bool = False) str¶
Dump a CDF file’s records in NASA’s own cdfirsdump (-full -nopage) text format, byte-for-byte (verified against real captures of that tool’s own output - see tests/nasa_compat_repr). Built on the same physical-order walk as for_each_record, not load()’s reconstructed view.
- Parameters:
- pathstr
Path to the CDF file.
- radixint, default 10
10 (decimal) or 16 (hex, “0x” + 16 uppercase digits) for record offsets.
- show_databool, default False
Hex-dump VVR/CVVR payload bytes (cdfirsdump’s -data).
- summarybool, default False
Append the closing record-type summary table (cdfirsdump’s default -summary behavior; this binding defaults to False to match this project’s own existing default, not NASA’s).
- Returns:
- str
The full dump text, ready to print or write to a file.
- pycdfpp.debug.nasa_compat_dump_from_offset(path: str, offset: SupportsInt, radix: SupportsInt = 10, show_data: bool = False) str¶
Same as nasa_compat_dump, but starts the walk at a given byte offset instead of the file start - no magic-number preamble is printed (matching a real cdfirsdump -offset capture).
- Parameters:
- pathstr
Path to the CDF file.
- offsetint
Byte offset to start the walk at.
- radixint, default 10
10 (decimal) or 16 (hex) for record offsets.
- show_databool, default False
Hex-dump VVR/CVVR payload bytes.
- Returns:
- str
The dump text from that offset onward.
- pycdfpp.debug.nasa_compat_dump_brief(path: str) str¶
Dump only the record-type summary table (cdfirsdump’s -brief, the real tool’s own default level), byte-for-byte, including its “(+16 if with checksum)” quirk which is always shown at brief level regardless of the file (see nasa_compat_repr.hpp’s own notes on why).
- Parameters:
- pathstr
Path to the CDF file.
- Returns:
- str
The banner + summary table text.