Skip to content

API reference

Generated from this package's own docstrings (numpydoc style -- see this repository's CLAUDE.md for the convention). This page covers pygdb's public surface -- everything listed in pygdb.__all__ -- grouped the way the quick start introduces them.

The GDB class

pygdb.GDB(path, include_unlisted_blobs=False)

High-level, name-based view of a single .gdb file.

Parameters:

Name Type Description Default
path str

Path to the .gdb file.

required
include_unlisted_blobs bool

A file that carries a blob directory (see Notes) lists exactly one live blob per (line, channel). By default a blob the directory does not list -- an all-zero slot -- is skipped, as if it were absent, and one GDBParseWarning says how many (line, channel) blobs in which channels were skipped. Pass True to read them anyway (the last copy in chain order). A file with no directory, or an all-zero one, is unaffected: every blob in the chain is served.

False

Raises:

Type Description
ValueError

At construction time, if path doesn't start with the expected .gdb magic -- unlike the module-level functions in gdb_reader/registry (which warn and return empty results), since a GDB object that isn't backed by a real .gdb file can't usefully do anything at all.

Examples:

>>> db = GDB("survey.gdb")
>>> db.line_names[:3]
['L1000', 'L1001', 'L1010']
>>> db.channels_on_line("L1000")[:3]
['Fiducial', 'Easting', 'Northing']
>>> db.read("L1000", "Easting")[:3]
[612345.6, 612346.1, 612346.7]
>>> db.compression.name
'DB_COMP_NONE'
>>> db.coordinate_systems
['NAD83 / UTM zone 11N', 'WGS 84']
Notes

Channel and line tables are read once, lazily, on first access, and cached; the (line, channel) -> blob index used by read() and channels_on_line() is likewise built once (a full blob-chain walk) on first use. Unlike the module-level gdb_reader functions this class is built on (which each reopen path fresh, for statelessness), GDB opens path once at construction and reuses that handle for every read()/iter_line() call -- reopening per call was measured at ~1.7-1.9x slower against this project's real sample corpus (see the Rust-plan's M4 notes). Close it (db.close(), or use GDB as a context manager) when done with it, or just let it get garbage-collected -- __del__ closes it too, as a safety net.

Which copy of a blob is read. The blob chain is append-only, so a (line, channel) can have several blobs, an older stale one and the current one, in either order (issue #2). The file's persisted blob directory (docs/spec.md section 2.2) says which is current, and it is the only thing consulted. A directory entry is used only if its start page lands on a walked blob header whose blob_index equals the slot and whose page count matches; if a non-zero entry fails that check, the last copy in chain order is used and a GDBParseWarning says so. Without a usable directory the last copy in chain order is used, with a warning if a pair is duplicated: that is a guess, and not always the current copy.

Source code in pygdb/gdb.py
def __init__(self, path: str, include_unlisted_blobs: bool = False):
    self.path = path
    self.include_unlisted_blobs = include_unlisted_blobs
    self._file = open(path, "rb")
    header = self._file.read(4096)
    if not check_magic(header):
        self._file.close()
        raise ValueError(
            f"{path}: does not start with the expected '!CBD' magic -- "
            f"not a recognized .gdb file"
        )
    self._fields = header_fields(header)
    unseen.check_header(path, header)
    unseen.check_admin_objects(path)
    self._channels: Optional[List[ChannelRecord]] = None
    self._lines: Optional[List[LineRecord]] = None
    self._lines_exact = False
    self._channels_by_name: Optional[Dict[str, List[ChannelRecord]]] = None
    self._lines_by_name: Optional[Dict[str, List[LineRecord]]] = None
    self._blob_index: Optional[Dict[Tuple[int, int], BlobHeader]] = None
    self._coordinate_systems: Optional[List[str]] = None
    self._coordinate_channels: Optional[Dict[str, Optional[str]]] = None
    self._channel_settings: Optional[Dict[str, Dict[str, str]]] = None
    self._projection_parameters: Optional[Dict[str, ProjectionParameters]] = None
    self._channel_makers: Optional[Dict[str, ChannelMaker]] = None
    self._display_lists: Optional[List[List[DisplayListEntry]]] = None

channel_makers property

dict of {str : ChannelMaker}: How each channel was made, from the MAKER records in this file's registry (docs/spec.md section 9) -- the tool ("newchan.gx", a math expression, "newxy.gx", ...), its label, and the parameters it ran with, e.g. a derived channel's formula. Only channels with a record are present; see pygdb.registry.find_channel_makers.

channel_names property

list of str: [c.name for c in self.channels].

channel_settings property

dict of {str : dict of {str : str}}: Real per-channel settings recorded in this file's own internal registry (docs/provenance/notes.md section 6.8c), e.g. {"raw_mag": {"UNITS": "nT"}, ...}. Only a channel with at least one populated registry key is present; most real files have none for most channels -- see pygdb.registry.find_channel_settings for what is and isn't decoded (the registry's key/value entries, not its nested MAKER records, which are channel_makers).

channels property

list of ChannelRecord: This file's channel table, read once and cached.

chans_max property

int or None: The file's channel-table capacity (header offset 24).

comp_level property

int or None: The file's declared compression level (header offset 120).

compression property

CompressionInfo: The file's declared compression mode.

coordinate_channels property

dict of {str : str or None}: Which real channel plays the X/Y/Z coordinate role, per this file's own internal registry (docs/provenance/notes.md section 6.8b) -- a directly-decodable alternative to guessing from channel-naming conventions. Always {"X": ..., "Y": ..., "Z": ...}; a role this file's registry doesn't confirm a real channel for (absent entirely, or a real but ambiguous/blank entry -- see pygdb.registry.find_channel_roles) is None, not omitted.

Confirmed present and correctly resolvable on every one of this project's 22 real sample files for X/Y (100%), 2 of 22 for Z -- to_geoh5 uses this as its coordinate-channel default, falling back to "Easting"/"Northing" only when a role isn't confirmed here.

coordinate_systems property

list of str: Best-effort coordinate-system/map-projection names found in this file's REG/IPJ administrative-blob content (docs/spec.md section 8-9). An empty list just means none were found -- not every real file has this content, and even when it does, this is a name-only extraction, not a full projection definition.

display_lists property

list of list of DisplayListEntry: The file's Display List objects (docs/spec.md section 9) -- [LIKELY] the channels shown in the database's spreadsheet view -- each entry resolved by handle to the channel's current name. See pygdb.registry.find_display_lists.

line_names property

list of str: [l.name for l in self.lines].

lines property

list of LineRecord: This file's line table, read once and cached.

page_size property

int or None: The file's page size in bytes (header offset 100).

projection_parameters property

dict of {str : ProjectionParameters}: Real geodetic parameters for each coordinate system named in coordinate_systems, decoded from this file's own internal IPJ registry (docs/spec.md section 8, docs/provenance/notes.md section 6.7b) -- datum, ellipsoid, and (when this coordinate system is a projected one) central meridian, scale factor, and false easting/northing. An empty dict just means no IPJ content was found, same as coordinate_systems; see pygdb.registry.find_projection_parameters for exactly what is and isn't decoded.

channel(name)

Look up a channel by name.

Parameters:

Name Type Description Default
name str or tuple of (str, int)

A plain channel name, or (name, occurrence) (occurrence a 0-based index into every channel sharing that name, in .channels order) to pick a specific one explicitly rather than relying on name alone being unambiguous.

required

Returns:

Type Description
ChannelRecord

The matching channel.

Raises:

Type Description
KeyError

If no channel has this name.

ValueError

If more than one channel shares this name and no occurrence was given -- a real, if unusual, on-disk possibility (confirmed for real on a sample file with two channels each named UTC, RADAR, and RAWMAG), for which there's no file-wide way to pick the "right" one without a line to disambiguate against. read() already disambiguates this automatically using line context (see _resolve_channel_on_line).

Source code in pygdb/gdb.py
def channel(self, name: ChannelRef) -> ChannelRecord:
    """
    Look up a channel by name.

    Parameters
    ----------
    name : str or tuple of (str, int)
        A plain channel name, or `(name, occurrence)`
        (`occurrence` a 0-based index into every channel sharing
        that name, in `.channels` order) to pick a specific one
        explicitly rather than relying on `name` alone being
        unambiguous.

    Returns
    -------
    ChannelRecord
        The matching channel.

    Raises
    ------
    KeyError
        If no channel has this name.
    ValueError
        If more than one channel shares this name and no
        `occurrence` was given -- a real, if unusual, on-disk
        possibility (confirmed for real on a sample file with two
        channels each named `UTC`, `RADAR`, and `RAWMAG`), for
        which there's no file-wide way to pick the "right" one
        without a line to disambiguate against. `read()` already
        disambiguates this automatically using line context (see
        `_resolve_channel_on_line`).
    """
    if self._channels_by_name is None:
        self.channels  # populate the cache
    if isinstance(name, tuple):
        actual_name, occurrence = name
        return self._nth_by_name(self._channels_by_name, actual_name, occurrence, "channel")
    matches = self._channels_by_name.get(name)
    if not matches:
        raise KeyError(f"{self.path}: no channel named {name!r}")
    if len(matches) > 1:
        raise ValueError(
            f"{self.path}: {len(matches)} channels are named {name!r} -- "
            f"ambiguous without a line to disambiguate against; use "
            f"read(line, name) (which resolves this using the line's "
            f"own data), pass (name, occurrence) to pick a specific "
            f"one explicitly, or pick a ChannelRecord from .channels "
            f"yourself"
        )
    return matches[0]

channels_on_line(line)

Names of channels that actually have data on line.

Parameters:

Name Type Description Default
line str or tuple or LineRecord

A line name, a (name, occurrence) pair (see line()), or a LineRecord.

required

Returns:

Type Description
list of str

Channel names with a real data blob recorded for line -- the format stores a sparse (line, channel) grid (docs/spec.md section 1), so most lines only populate a subset of this file's full channel list. If two channels share a name and both have data on this line, that name appears twice here (a list, so nothing is silently dropped) -- use iter_line() instead if you need the actual ChannelRecord for each entry, not just its name.

Source code in pygdb/gdb.py
def channels_on_line(self, line: LineRef) -> List[str]:
    """
    Names of channels that actually have data on `line`.

    Parameters
    ----------
    line : str or tuple or LineRecord
        A line name, a `(name, occurrence)` pair (see `line()`),
        or a `LineRecord`.

    Returns
    -------
    list of str
        Channel names with a real data blob recorded for `line` --
        the format stores a sparse (line, channel) grid
        (docs/spec.md section 1), so most lines only populate a
        subset of this file's full channel list. If two channels
        share a name and both have data on this line, that name
        appears twice here (a list, so nothing is silently
        dropped) -- use `iter_line()` instead if you need the
        actual `ChannelRecord` for each entry, not just its name.
    """
    line_rec = self._resolve_line(line)
    return [c.name for c, _blob in self._channels_with_data_on_line(line_rec)]

close()

Close the underlying file handle. Safe to call more than once.

Source code in pygdb/gdb.py
def close(self) -> None:
    """Close the underlying file handle. Safe to call more than once."""
    self._file.close()

iter_line(line)

Iterate every channel's data on one line.

Parameters:

Name Type Description Default
line str or tuple or LineRecord

A line name, a (name, occurrence) pair, or a LineRecord.

required

Yields:

Name Type Description
channel ChannelRecord

The channel itself, not just its name -- deliberately, so that if two channels share a name and both have data on this line, both still come through as distinct, fully-identified entries. dict(db.iter_line(line)) keyed by the records themselves preserves that; collapsing to channel.name yourself reintroduces the same collision read() raises on, so do that deliberately if you do it at all.

values ndarray

Same as read(line, channel) would give for that exact channel.

Notes

Uses _channels_with_data_on_line directly rather than one read() call per channel (saving the repeated name lookups).

A rayon-based parallel batch decoder was tried for this and removed: benchmarked against this project's real sample corpus, it was consistently ~3x slower than plain sequential calls at every scale tried, since this format's chunks are small enough that the Rust decoder (pygdb._native) already finishes each one in a fraction of a millisecond -- not enough work per chunk to amortize rayon's per-task dispatch cost. See rust/src/lib.rs's module doc for the full note.

Source code in pygdb/gdb.py
def iter_line(self, line: LineRef) -> Iterator[Tuple[ChannelRecord, np.ndarray]]:
    """
    Iterate every channel's data on one line.

    Parameters
    ----------
    line : str or tuple or LineRecord
        A line name, a `(name, occurrence)` pair, or a `LineRecord`.

    Yields
    ------
    channel : ChannelRecord
        The channel itself, not just its name -- deliberately, so
        that if two channels share a name and both have data on
        this line, both still come through as distinct,
        fully-identified entries. `dict(db.iter_line(line))` keyed
        by the records themselves preserves that; collapsing to
        `channel.name` yourself reintroduces the same collision
        `read()` raises on, so do that deliberately if you do it
        at all.
    values : numpy.ndarray
        Same as `read(line, channel)` would give for that exact
        channel.

    Notes
    -----
    Uses `_channels_with_data_on_line` directly rather than one
    `read()` call per channel (saving the repeated name lookups).

    A `rayon`-based parallel batch decoder was tried for this and
    removed: benchmarked against this project's real sample
    corpus, it was consistently ~3x *slower* than plain sequential
    calls at every scale tried, since this format's chunks are
    small enough that the Rust decoder (`pygdb._native`) already
    finishes each one in a fraction of a millisecond -- not enough
    work per chunk to amortize rayon's per-task dispatch cost. See
    `rust/src/lib.rs`'s module doc for the full note.
    """
    line_rec = self._resolve_line(line)
    for c, blob in self._channels_with_data_on_line(line_rec):
        values = read_blob_values(
            self.path, blob, c,
            comp_level=self.comp_level or 0, page_size=self.page_size,
            file=self._file,
        )
        yield c, values

line(name)

Look up a line by name.

Parameters:

Name Type Description Default
name str or tuple of (str, int)

A plain line name, or (name, occurrence) (occurrence a 0-based index into every line sharing that name, in .lines order) to pick a specific one explicitly.

required

Returns:

Type Description
LineRecord

The matching line.

Raises:

Type Description
KeyError

If no line has this name.

ValueError

If more than one line shares this name and no occurrence was given -- the line table has the same on-disk shape as the channel table (see channel()), with nothing in the format forbidding a duplicate name there either; not yet observed on a real file, but handled the same way on principle rather than left as a silent last-one-wins lookup.

Source code in pygdb/gdb.py
def line(self, name: LineRef) -> LineRecord:
    """
    Look up a line by name.

    Parameters
    ----------
    name : str or tuple of (str, int)
        A plain line name, or `(name, occurrence)` (`occurrence` a
        0-based index into every line sharing that name, in
        `.lines` order) to pick a specific one explicitly.

    Returns
    -------
    LineRecord
        The matching line.

    Raises
    ------
    KeyError
        If no line has this name.
    ValueError
        If more than one line shares this name and no `occurrence`
        was given -- the line table has the same on-disk shape as
        the channel table (see `channel()`), with nothing in the
        format forbidding a duplicate name there either; not yet
        observed on a real file, but handled the same way on
        principle rather than left as a silent last-one-wins
        lookup.
    """
    if self._lines_by_name is None:
        self.lines  # populate the cache
    if isinstance(name, tuple):
        actual_name, occurrence = name
        return self._nth_by_name(self._lines_by_name, actual_name, occurrence, "line")
    matches = self._lines_by_name.get(name)
    if not matches:
        raise KeyError(f"{self.path}: no line named {name!r}")
    if len(matches) > 1:
        raise ValueError(
            f"{self.path}: {len(matches)} lines are named {name!r} -- "
            f"ambiguous; pass (name, occurrence) to pick a specific "
            f"one explicitly, or pick a LineRecord from .lines yourself"
        )
    return matches[0]

read(line, channel)

Random access by name: decode every value recorded for channel on line.

Parameters:

Name Type Description Default
line str or tuple or LineRecord

A line name, a (name, occurrence) pair, or a LineRecord.

required
channel str or tuple or ChannelRecord

A channel name, a (name, occurrence) pair, or a ChannelRecord.

required

Returns:

Type Description
ndarray

1-D for an ordinary scalar channel, 2-D (n_rows, channel.array_width) for a VA/array channel (docs/spec.md section 5), dtype matching the channel's GS_* type, or a fixed-width Unicode dtype for a string-typed channel.

Raises:

Type Description
KeyError

If line or channel isn't a name this file has.

ValueError

If channel is a plain name shared by more than one channel (see channel()) and it's still ambiguous after resolving it using line's own data (_resolve_channel_on_line) -- i.e. more than one same-named channel has data on this exact line. Pass (name, occurrence) or a specific ChannelRecord to sidestep either lookup.

Warns:

Type Description
GDBParseWarning

Per read_blob_values, if line/channel are valid names but this specific (line, channel) pair has no data blob, or its data can't be decoded -- an empty array is returned rather than raising, consistent with the rest of this package's degrade-gracefully philosophy for decode-time problems, as opposed to a plain lookup-by-name mistake (which does raise).

Source code in pygdb/gdb.py
def read(self, line: LineRef, channel: ChannelRef) -> np.ndarray:
    """
    Random access by name: decode every value recorded for `channel` on `line`.

    Parameters
    ----------
    line : str or tuple or LineRecord
        A line name, a `(name, occurrence)` pair, or a `LineRecord`.
    channel : str or tuple or ChannelRecord
        A channel name, a `(name, occurrence)` pair, or a
        `ChannelRecord`.

    Returns
    -------
    numpy.ndarray
        1-D for an ordinary scalar channel, 2-D `(n_rows,
        channel.array_width)` for a VA/array channel (docs/spec.md
        section 5), dtype matching the channel's `GS_*` type, or a
        fixed-width Unicode dtype for a string-typed channel.

    Raises
    ------
    KeyError
        If `line` or `channel` isn't a name this file has.
    ValueError
        If `channel` is a plain name shared by more than one
        channel (see `channel()`) and it's *still* ambiguous after
        resolving it using `line`'s own data
        (`_resolve_channel_on_line`) -- i.e. more than one
        same-named channel has data on this exact line. Pass
        `(name, occurrence)` or a specific `ChannelRecord` to
        sidestep either lookup.

    Warns
    -----
    GDBParseWarning
        Per `read_blob_values`, if `line`/`channel` are valid
        names but this specific (line, channel) pair has no data
        blob, or its data can't be decoded -- an empty array is
        returned rather than raising, consistent with the rest of
        this package's degrade-gracefully philosophy for
        decode-time problems, as opposed to a plain lookup-by-name
        mistake (which does raise).
    """
    line_rec = self._resolve_line(line)
    if isinstance(channel, ChannelRecord):
        chan_rec = channel
        blob = self._ensure_blob_index().get((line_rec.index, chan_rec.index))
    elif isinstance(channel, tuple):
        chan_rec = self.channel(channel)
        blob = self._ensure_blob_index().get((line_rec.index, chan_rec.index))
    else:
        chan_rec, blob = self._resolve_channel_on_line(line_rec, channel)
    if blob is None:
        warnings.warn(
            f"{self.path}: no data blob for line {line_rec.name!r}, "
            f"channel {chan_rec.name!r} -- this (line, channel) pair "
            f"was likely never recorded (the format's grid is sparse, "
            f"docs/spec.md section 1)",
            GDBParseWarning, stacklevel=2,
        )
        return np.array([])
    return read_blob_values(
        self.path, blob, chan_rec,
        comp_level=self.comp_level or 0, page_size=self.page_size,
        file=self._file,
    )

to_dataframe(line=None)

Build a pandas.DataFrame -- one row per station.

Needs the optional pandas dependency (pip install python-gdb[pandas]), imported lazily here so importing pygdb itself never requires it.

Parameters:

Name Type Description Default
line str, int, or LineRecord

If given, scope the export to that one line only, mirroring to_xarray(line)'s exact scope. df.attrs gets line_name, line_category (line_rec.category_name), and path -- the same three keys to_xarray's Dataset.attrs carries -- rather than a "line" column, since every row already belongs to the one given line.

If omitted (the default), export every line with real data, concatenated into one table -- "line"/ "line_category" columns are added so rows from different lines stay distinguishable (pandas.DataFrame.attrs is a single, frame-level dict, not one per line, so it can't hold this the way single-line mode's df.attrs does). See Warns for what happens if a real channel collides with one of these two reserved names.

None

Returns:

Type Description
DataFrame

Raises:

Type Description
ImportError

If the optional pandas dependency isn't installed.

Warns:

Type Description
GDBParseWarning

In whole-file mode, if a real channel collides with the reserved "line"/"line_category" column names -- confirmed real, not hypothetical: a real Ontario sample file has a channel literally named "line" -- in which case the channel's own column is renamed f"channel_{name}" rather than silently overwriting the reserved column every whole-file caller relies on to tell rows apart. Also raised for a duplicate channel name (see Notes) and a row-count mismatch (see Notes).

Notes

A VA/array channel (docs/spec.md section 5) is exported as one column per element, f"{name}[{j}]" for j in range(array_width) -- the same flattening to_geoh5 uses, for the same reason: there's no natural "one cell holds an array" representation in a plain 2D table.

If two channels share a name and both have data on a line, the second (and any later) occurrence's column is disambiguated as f"{name}[{occurrence}]", identical numbering to to_xarray/ to_geoh5 (see _disambiguate_names).

If channels on one line don't all decode to the same row count (a truncated/corrupt file), the short channel's column is padded with missing values out to the max row count across that line's channels, rather than being dropped or forcing other channels to truncate to match -- see _pad_to_length's docstring for why padding, specifically, is the right default here where it wasn't for to_xarray/to_geoh5. A channel entirely absent on one line in whole-file mode (present on some lines, not others -- a normal, sparse case, docs/spec.md section 1) needs no special handling here: it's simply missing from that line's own per-line frame, and pandas.concat's ordinary union-of-columns behavior fills the gap with missing values across the whole table -- no warning, since a channel simply not being on a line is the format's normal baseline, not a sign of trouble.

Source code in pygdb/gdb.py
def to_dataframe(self, line: Optional[LineRef] = None) -> "pd.DataFrame":
    """
    Build a `pandas.DataFrame` -- one row per station.

    Needs the optional `pandas` dependency (`pip install
    python-gdb[pandas]`), imported lazily here so importing
    `pygdb` itself never requires it.

    Parameters
    ----------
    line : str, int, or LineRecord, optional
        If given, scope the export to that one line only,
        mirroring `to_xarray(line)`'s exact scope. `df.attrs` gets
        `line_name`, `line_category` (`line_rec.category_name`),
        and `path` -- the same three keys `to_xarray`'s
        `Dataset.attrs` carries -- rather than a `"line"` column,
        since every row already belongs to the one given line.

        If omitted (the default), export every line with real
        data, concatenated into one table -- `"line"`/
        `"line_category"` columns are added so rows from different
        lines stay distinguishable (`pandas.DataFrame.attrs` is a
        single, frame-level dict, not one per line, so it can't
        hold this the way single-line mode's `df.attrs` does). See
        Warns for what happens if a real channel collides with one
        of these two reserved names.

    Returns
    -------
    pandas.DataFrame

    Raises
    ------
    ImportError
        If the optional `pandas` dependency isn't installed.

    Warns
    -----
    GDBParseWarning
        In whole-file mode, if a real channel collides with the
        reserved `"line"`/`"line_category"` column names --
        confirmed real, not hypothetical: a real Ontario sample
        file has a channel literally named `"line"` -- in which
        case the *channel's* own column is renamed
        `f"channel_{name}"` rather than silently overwriting the
        reserved column every whole-file caller relies on to tell
        rows apart. Also raised for a duplicate channel name (see
        Notes) and a row-count mismatch (see Notes).

    Notes
    -----
    A VA/array channel (docs/spec.md section 5) is exported as one
    column per element, `f"{name}[{j}]"` for `j` in
    `range(array_width)` -- the same flattening `to_geoh5` uses,
    for the same reason: there's no natural "one cell holds an
    array" representation in a plain 2D table.

    If two channels share a name and both have data on a line, the
    second (and any later) occurrence's column is disambiguated as
    `f"{name}[{occurrence}]"`, identical numbering to `to_xarray`/
    `to_geoh5` (see `_disambiguate_names`).

    If channels on one line don't all decode to the same row count
    (a truncated/corrupt file), the short channel's column is
    padded with missing values out to the max row count across
    that line's channels, rather than being dropped or forcing
    other channels to truncate to match -- see `_pad_to_length`'s
    docstring for why padding, specifically, is the right default
    here where it wasn't for `to_xarray`/`to_geoh5`. A channel
    entirely absent on one line in whole-file mode (present on
    some lines, not others -- a normal, sparse case, docs/spec.md
    section 1) needs no special handling here: it's simply missing
    from that line's own per-line frame, and `pandas.concat`'s
    ordinary union-of-columns behavior fills the gap with missing
    values across the whole table -- no warning, since a channel
    simply not being on a line is the format's normal baseline,
    not a sign of trouble.
    """
    try:
        import pandas as pd
    except ImportError as e:
        raise ImportError(
            "to_dataframe() needs the optional 'pandas' dependency -- "
            "install with `pip install python-gdb[pandas]`"
        ) from e

    whole_file = line is None
    line_recs = self.lines if whole_file else [self._resolve_line(line)]

    frames = []
    for line_rec in line_recs:
        pairs = self._channels_with_data_on_line(line_rec)
        if not pairs:
            continue

        decoded = [
            (c, read_blob_values(
                self.path, blob, c,
                comp_level=self.comp_level or 0, page_size=self.page_size,
                file=self._file,
            ))
            for c, blob in pairs
        ]
        row_count = max((len(values) for _c, values in decoded), default=0)

        columns: Dict[str, object] = {}
        if whole_file:
            columns["line"] = [line_rec.name] * row_count
            columns["line_category"] = [line_rec.category_name] * row_count

        for var_name, c, values in self._disambiguate_names(line_rec, decoded):
            if var_name in columns:
                # A real channel can collide with "line"/"line_category"
                # (confirmed real: a real Ontario file has a channel
                # literally named "line") -- caught here generically,
                # not just for those two names, since any future
                # reserved column would hit the same hazard. Rename the
                # channel's own column rather than silently letting it
                # overwrite the reserved one.
                original = var_name
                var_name = f"channel_{var_name}"
                warnings.warn(
                    f"{self.path}: line {line_rec.name!r} has a channel "
                    f"named {original!r}, which collides with a column "
                    f"name this method already uses -- using {var_name!r} "
                    f"for this channel's own data instead",
                    GDBParseWarning, stacklevel=2,
                )
            if c.is_array:
                for j in range(c.array_width):
                    columns[f"{var_name}[{j}]"] = self._pad_to_length(
                        values[:, j], row_count, line_rec, c, f"{var_name}[{j}]", pd,
                    )
            else:
                columns[var_name] = self._pad_to_length(
                    values, row_count, line_rec, c, var_name, pd,
                )
        frames.append(pd.DataFrame(columns))

    if not frames:
        return pd.DataFrame()

    df = pd.concat(frames, ignore_index=True)
    if not whole_file:
        line_rec = line_recs[0]
        df.attrs = {
            "line_name": line_rec.name,
            "line_category": line_rec.category_name,
            "path": self.path,
        }
    return df

to_geoh5(path, *, x_channel=None, y_channel=None, z_channel=None)

Export every line's data to a new .geoh5 file at path.

One Points object per line (holding real per-line data, i.e. it has at least one channel with data), grouped under one ContainerGroup named after this .gdb file, inside a geoh5py.Workspace. Needs the optional geoh5py dependency (pip install python-gdb[geoh5]), imported lazily here so importing pygdb itself never requires it.

Unlike to_xarray (one line, in memory, no geometry needed), this is whole-file and writes directly to disk, since a geoh5py.Workspace is inherently file-backed and .geoh5's own natural unit is one file holding a whole survey's worth of named objects, not one line at a time.

Parameters:

Name Type Description Default
path str

Path to the .geoh5 file to create (or append to -- opened with Workspace(path, mode="a")).

required
x_channel str

Name the channels providing each line's Points.vertices. Left unset (None, the default for all three), this first tries self.coordinate_channels -- this file's own internal registry of which real channel plays which coordinate role (docs/provenance/notes.md section 6.8b), directly decoded, not guessed -- confirmed present and correct on 100% of this project's real sample corpus for X/Y. Only when that registry doesn't confirm a role does this fall back to the literal names "Easting"/ "Northing" (X/Y) or no channel at all (Z, meaning every vertex gets Z = 0.0). Passing an explicit channel name always wins over both.

None
y_channel str

Name the channels providing each line's Points.vertices. Left unset (None, the default for all three), this first tries self.coordinate_channels -- this file's own internal registry of which real channel plays which coordinate role (docs/provenance/notes.md section 6.8b), directly decoded, not guessed -- confirmed present and correct on 100% of this project's real sample corpus for X/Y. Only when that registry doesn't confirm a role does this fall back to the literal names "Easting"/ "Northing" (X/Y) or no channel at all (Z, meaning every vertex gets Z = 0.0). Passing an explicit channel name always wins over both.

None
z_channel str

Name the channels providing each line's Points.vertices. Left unset (None, the default for all three), this first tries self.coordinate_channels -- this file's own internal registry of which real channel plays which coordinate role (docs/provenance/notes.md section 6.8b), directly decoded, not guessed -- confirmed present and correct on 100% of this project's real sample corpus for X/Y. Only when that registry doesn't confirm a role does this fall back to the literal names "Easting"/ "Northing" (X/Y) or no channel at all (Z, meaning every vertex gets Z = 0.0). Passing an explicit channel name always wins over both.

None

Raises:

Type Description
ImportError

If the optional geoh5py dependency isn't installed.

Warns:

Type Description
GDBParseWarning

A line missing its resolved X or Y channel is skipped entirely (no Points object created for it) rather than guessed at or defaulted to (0, 0) -- this reader never guesses what a channel means when it can't confirm one (see to_xarray's docstring); a missing Z, by contrast, is normal and never blocks export. Also raised for a duplicate channel name (see Notes) and a row-count mismatch (see Notes).

Notes

Array/VA channels (docs/spec.md section 5): .geoh5 (per geoh5py, checked directly against its real data/ module source) has no Data type holding more than one value per vertex -- the plain numeric types silently ravel() anything with ndim > 1. So each array channel is exported as one Data entry per column, named f"{name}[{j}]" for j in range(array_width), tied back together with a PropertyGroup named after the channel (ObjectBase.add_data(..., property_group=name)) -- the same "one Data per gate/column, grouped" pattern geoh5py's own built-in survey types (AirborneTEMSurvey et al.) use internally for multi-gate EM decay-curve data, not a workaround invented here.

Duplicate channel names: disambiguated exactly like to_xarray (see _disambiguate_names) -- f"{name}[{occurrence}]" for every occurrence after the first, same numbering as channel(("name", occurrence)).

Row-count mismatches: .geoh5 Data with VERTEX association must match the parent object's vertex count exactly -- there's no analogue to to_xarray's per-channel dimension escape hatch. A channel that decodes to a different row count than this line's vertex count (from x_channel) is skipped (that channel only, not the whole line).

Not attempted: coordinate-system/CRS export -- self.coordinate_systems only returns best-effort names (no EPSG codes or full projection definitions), which isn't enough to populate .geoh5's real CRS metadata correctly.

Source code in pygdb/gdb.py
def to_geoh5(
    self,
    path: str,
    *,
    x_channel: Optional[str] = None,
    y_channel: Optional[str] = None,
    z_channel: Optional[str] = None,
) -> None:
    """
    Export every line's data to a new `.geoh5` file at `path`.

    One `Points` object per line (holding real per-line data,
    i.e. it has at least one channel with data), grouped under one
    `ContainerGroup` named after this `.gdb` file, inside a
    `geoh5py.Workspace`. Needs the optional `geoh5py` dependency
    (`pip install python-gdb[geoh5]`), imported lazily here so
    importing `pygdb` itself never requires it.

    Unlike `to_xarray` (one line, in memory, no geometry needed),
    this is whole-file and writes directly to disk, since a
    `geoh5py.Workspace` is inherently file-backed and `.geoh5`'s own
    natural unit is one file holding a whole survey's worth of named
    objects, not one line at a time.

    Parameters
    ----------
    path : str
        Path to the `.geoh5` file to create (or append to --
        opened with `Workspace(path, mode="a")`).
    x_channel, y_channel, z_channel : str, optional
        Name the channels providing each line's `Points.vertices`.
        Left unset (`None`, the default for all three), this first
        tries `self.coordinate_channels` -- this file's own
        internal registry of which real channel plays which
        coordinate role (docs/provenance/notes.md section 6.8b),
        directly decoded, not guessed -- confirmed present and
        correct on 100% of this project's real sample corpus for
        X/Y. Only when that registry doesn't confirm a role does
        this fall back to the literal names `"Easting"`/
        `"Northing"` (X/Y) or no channel at all (Z, meaning every
        vertex gets `Z = 0.0`). Passing an explicit channel name
        always wins over both.

    Raises
    ------
    ImportError
        If the optional `geoh5py` dependency isn't installed.

    Warns
    -----
    GDBParseWarning
        A line missing its resolved X or Y channel is **skipped
        entirely** (no `Points` object created for it) rather than
        guessed at or defaulted to `(0, 0)` -- this reader never
        guesses what a channel means when it can't confirm one
        (see `to_xarray`'s docstring); a missing Z, by contrast, is
        normal and never blocks export. Also raised for a
        duplicate channel name (see Notes) and a row-count
        mismatch (see Notes).

    Notes
    -----
    **Array/VA channels** (docs/spec.md section 5): `.geoh5` (per
    `geoh5py`, checked directly against its real `data/` module
    source) has no `Data` type holding more than one value per
    vertex -- the plain numeric types silently `ravel()` anything
    with `ndim > 1`. So each array channel is exported as **one
    `Data` entry per column**, named `f"{name}[{j}]"` for `j` in
    `range(array_width)`, tied back together with a `PropertyGroup`
    named after the channel (`ObjectBase.add_data(...,
    property_group=name)`) -- the same "one `Data` per gate/column,
    grouped" pattern `geoh5py`'s own built-in survey types
    (`AirborneTEMSurvey` et al.) use internally for multi-gate EM
    decay-curve data, not a workaround invented here.

    **Duplicate channel names**: disambiguated exactly like
    `to_xarray` (see `_disambiguate_names`) -- `f"{name}[{occurrence}]"`
    for every occurrence after the first, same numbering as
    `channel(("name", occurrence))`.

    **Row-count mismatches**: `.geoh5` `Data` with `VERTEX`
    association must match the parent object's vertex count
    exactly -- there's no analogue to `to_xarray`'s per-channel
    dimension escape hatch. A channel that decodes to a different
    row count than this line's vertex count (from `x_channel`) is
    **skipped** (that channel only, not the whole line).

    **Not attempted**: coordinate-system/CRS export --
    `self.coordinate_systems` only returns best-effort names (no
    EPSG codes or full projection definitions), which isn't enough
    to populate `.geoh5`'s real CRS metadata correctly.
    """
    try:
        import geoh5py  # noqa: F401 -- import-only check, see below
    except ImportError as e:
        raise ImportError(
            "to_geoh5() needs the optional 'geoh5py' dependency -- "
            "install with `pip install python-gdb[geoh5]`"
        ) from e
    # A bare `import geoh5py` (rather than importing these submodules
    # directly inside the `try`) is what makes the dependency check
    # actually fire in every case, including when `geoh5py`'s own
    # submodules are already cached in `sys.modules` from an earlier
    # import elsewhere in the process: a `from geoh5py.groups import
    # ContainerGroup` reuses an already-imported `geoh5py.groups`
    # without re-checking `geoh5py` itself, so testing for a missing
    # dependency by monkeypatching `sys.modules["geoh5py"] = None`
    # would otherwise silently not trigger this except block.
    from geoh5py.groups import ContainerGroup
    from geoh5py.objects import Points
    from geoh5py.workspace import Workspace

    if x_channel is None or y_channel is None or z_channel is None:
        detected = self.coordinate_channels
        if x_channel is None:
            x_channel = detected["X"] or "Easting"
        if y_channel is None:
            y_channel = detected["Y"] or "Northing"
        if z_channel is None:
            z_channel = detected["Z"]  # stays None (all-zero Z) if unconfirmed

    with Workspace(path, mode="a") as ws:
        file_group = ws.create_entity(ContainerGroup, entity={"name": Path(self.path).stem})

        for line_rec in self.lines:
            pairs = self._channels_with_data_on_line(line_rec)
            if not pairs:
                continue

            decoded = [
                (c, read_blob_values(
                    self.path, blob, c,
                    comp_level=self.comp_level or 0, page_size=self.page_size,
                    file=self._file,
                ))
                for c, blob in pairs
            ]

            def _find(name):
                for c, values in decoded:
                    if c.name == name:
                        return c, values
                return None, None

            x_chan, x_values = _find(x_channel)
            y_chan, y_values = _find(y_channel)
            if x_chan is None or y_chan is None:
                missing = [n for n, c in ((x_channel, x_chan), (y_channel, y_chan)) if c is None]
                warnings.warn(
                    f"{self.path}: line {line_rec.name!r} has no data for "
                    f"{missing!r} -- can't place its points; skipping this "
                    f"line entirely (pass x_channel=/y_channel= if this file "
                    f"uses different coordinate channel names)",
                    GDBParseWarning, stacklevel=2,
                )
                continue

            n_vertices = len(x_values)
            if len(y_values) != n_vertices:
                warnings.warn(
                    f"{self.path}: line {line_rec.name!r}: {x_channel!r} decoded "
                    f"{n_vertices} row(s) but {y_channel!r} decoded "
                    f"{len(y_values)} -- can't build consistent vertices; "
                    f"skipping this line entirely",
                    GDBParseWarning, stacklevel=2,
                )
                continue

            geometry_channels = {x_chan, y_chan}
            z_values = np.zeros(n_vertices)
            if z_channel is not None:
                z_chan, found_z_values = _find(z_channel)
                if z_chan is not None and len(found_z_values) == n_vertices:
                    z_values = found_z_values
                    geometry_channels.add(z_chan)
                elif z_chan is not None:
                    warnings.warn(
                        f"{self.path}: line {line_rec.name!r}: z_channel "
                        f"{z_channel!r} decoded {len(found_z_values)} row(s), "
                        f"expected {n_vertices} -- using all-zero Z for this "
                        f"line instead",
                        GDBParseWarning, stacklevel=2,
                    )

            vertices = np.column_stack(
                [x_values, y_values, z_values]
            ).astype(float)
            points_obj = ws.create_entity(
                Points,
                entity={"name": line_rec.name, "parent": file_group, "vertices": vertices},
            )
            points_obj.add_data({
                "line_category": {"values": line_rec.category_name, "association": "OBJECT"},
            })

            for var_name, c, values in self._disambiguate_names(line_rec, decoded):
                if c in geometry_channels:
                    continue
                if var_name == "line_category":
                    # A real channel can collide with the object-
                    # level metadata key added above (confirmed
                    # real risk -- see the "line"/"line_category"
                    # collisions already found in to_dataframe/
                    # to_xarray). geoh5py's own `add_data` already
                    # auto-renames on a name clash (so this isn't a
                    # data-loss risk the way it was for a plain
                    # dict/Dataset), but it does so silently -- warn
                    # explicitly here too, for the same reason this
                    # reader warns everywhere else a name had to
                    # change unexpectedly.
                    original = var_name
                    var_name = f"channel_{var_name}"
                    warnings.warn(
                        f"{self.path}: line {line_rec.name!r} channel "
                        f"{original!r} collides with the object-level "
                        f"metadata key this method always adds -- using "
                        f"{var_name!r} for this channel's own data instead",
                        GDBParseWarning, stacklevel=2,
                    )
                if len(values) != n_vertices:
                    warnings.warn(
                        f"{self.path}: line {line_rec.name!r} channel {c.name!r} "
                        f"decoded {len(values)} row(s), expected {n_vertices} "
                        f"(this line's vertex count) -- geoh5 data must match "
                        f"the object's vertex count exactly; skipping this "
                        f"channel",
                        GDBParseWarning, stacklevel=2,
                    )
                    continue

                description = f"{c.type_name} / {c.format_name}"
                if c.is_array:
                    description += f" / {c.array_basetype_name}"
                    points_obj.add_data(
                        {
                            f"{var_name}[{j}]": {
                                "values": values[:, j],
                                "association": "VERTEX",
                                "description": description,
                            }
                            for j in range(c.array_width)
                        },
                        property_group=var_name,
                    )
                else:
                    points_obj.add_data({
                        var_name: {
                            "values": values,
                            "association": "VERTEX",
                            "description": description,
                        },
                    })

to_xarray(line=None)

Build an xarray.Dataset.

Needs the optional xarray dependency (pip install python-gdb[xarray]), imported lazily here so importing pygdb itself never requires it.

Parameters:

Name Type Description Default
line str, int, or LineRecord

If given, scope the export to that one line only -- one data variable per channel that has data on that line, sharing a common "station" dimension. See _to_xarray_one_line's docstring for the full single-line behavior (array-channel _bin dimensions, the _station dimension-splitting escape hatch for a row-count mismatch, line_name/line_category attrs).

If omitted (the default), export every line with real data, stacked along a new "line" dimension into one Dataset -- see _to_xarray_whole_file's docstring for how row-count mismatches and channels missing on some lines are filled (NaN/""/the channel's own Geosoft dummy value, recorded as each variable's _FillValue attr) rather than each getting its own dimension the way single-line mode does.

None

Returns:

Type Description
Dataset

Raises:

Type Description
ImportError

If the optional xarray dependency isn't installed.

Notes

Either way, a channel name shared by more than one channel in this file's table is disambiguated as f"{name}[{occurrence}]" the same way, resolved file-wide (not per line) so a given channel's variable name never depends on which line(s) are being exported -- see _disambiguate_names.

Source code in pygdb/gdb.py
def to_xarray(self, line: Optional[LineRef] = None) -> "xr.Dataset":
    """
    Build an `xarray.Dataset`.

    Needs the optional `xarray` dependency (`pip install
    python-gdb[xarray]`), imported lazily here so importing
    `pygdb` itself never requires it.

    Parameters
    ----------
    line : str, int, or LineRecord, optional
        If given, scope the export to that one line only -- one
        data variable per channel that has data on that line,
        sharing a common `"station"` dimension. See
        `_to_xarray_one_line`'s docstring for the full single-line
        behavior (array-channel `_bin` dimensions, the `_station`
        dimension-splitting escape hatch for a row-count mismatch,
        `line_name`/`line_category` attrs).

        If omitted (the default), export every line with real
        data, stacked along a new `"line"` dimension into one
        `Dataset` -- see `_to_xarray_whole_file`'s docstring for
        how row-count mismatches and channels missing on some
        lines are filled (`NaN`/`""`/the channel's own Geosoft
        dummy value, recorded as each variable's `_FillValue`
        attr) rather than each getting its own dimension the way
        single-line mode does.

    Returns
    -------
    xarray.Dataset

    Raises
    ------
    ImportError
        If the optional `xarray` dependency isn't installed.

    Notes
    -----
    Either way, a channel name shared by more than one channel in
    this file's table is disambiguated as `f"{name}[{occurrence}]"`
    the same way, resolved **file-wide** (not per line) so a given
    channel's variable name never depends on which line(s) are
    being exported -- see `_disambiguate_names`.
    """
    try:
        import xarray as xr
    except ImportError as e:
        raise ImportError(
            "to_xarray() needs the optional 'xarray' dependency -- "
            "install with `pip install python-gdb[xarray]`"
        ) from e

    if line is not None:
        return self._to_xarray_one_line(self._resolve_line(line), xr)
    return self._to_xarray_whole_file(xr)

pygdb.CompressionInfo(code, name, codec) dataclass

This file's declared compression mode.

Header offset 120, docs/spec.md section 7.

Attributes:

Name Type Description
code int or None

Raw compression-level int32 (0, 1, or 2).

name str

The matching DB_COMP_* name.

codec str

The matching real codec name ("none", "lzrw1", or "zlib").

Notes

This describes what the file was configured with, not a guarantee every blob actually used it -- some real files declare DB_COMP_SPEED/DB_COMP_SIZE but contain zero compressed blobs (docs/spec.md section 7.6), and individual "bare" blobs inside a genuinely-compressed file can skip compression entirely (docs/spec.md section 7.4) -- read_blob_values() already detects and handles both cases automatically per-blob.

Records

pygdb.ChannelRecord(index, offset, name, dtype_code, format_code, raw, array_width=1, array_basetype_code=0, name_is_clean=True) dataclass

One record from a .gdb file's channel symbol table.

Attributes:

Name Type Description
index int

Slot index into the channel table.

offset int

Byte offset of this record within the file.

name str

Channel name.

dtype_code int

Raw int16 value: positive = GS_* type (see GS_TYPE_NUMPY_DTYPE), negative = -string_width.

format_code int

Raw int16 display-format code (see DB_CHAN_FORMAT_NAMES).

raw bytes

The record's raw, undecoded bytes.

array_width int, default 1

[CONFIRMED] relative offset +118, int16. 1 = plain scalar channel (the overwhelming majority of real channels seen). >1 = a true VA/array channel storing that many elements per fiducial "cell" -- e.g. 24 (time-decay gates) or 30 (depth layers) in the real AG106386 Georgetown conductivity file. Independently cross-checked against that same file's plain-text ASCII sibling (.dfn) format, which spells out "30F10.4" (Fortran-style: 30 repetitions of a float field) for the exact same channel name -- see docs/provenance/notes.md.

array_basetype_code int, default 0

[LIKELY] relative offset +86, int16.

name_is_clean bool, default True

False if the name field was NUL-unterminated or had non-printable bytes -- see docs/provenance/notes.md re: older (pre-2020, e.g. 1990s GEOTEM) files sometimes leaving unused capacity slots un-zeroed rather than clean, unlike the 2020 USGS samples.

Notes

eq=False keeps the default identity-based __eq__/__hash__ instead of dataclass's usual field-by-field one, so instances stay hashable (GDB.iter_line() yields these and documents dict(...) keyed by the record itself as safe -- see its docstring -- which needs __hash__ to actually work). Value equality between two separately-constructed-but-identical records is never used anywhere in this codebase; every real lookup returns the same cached instance from GDB.channels, so identity is all that's ever needed in practice.

array_basetype_name property

str: Human-readable array base-type name (a DB_ARRAY_BASETYPE_* name).

format_name property

str: Human-readable display-format name (a DB_CHAN_FORMAT_* name).

is_array property

bool: True for a real VA/array channel (array_width > 1). [CONFIRMED].

is_string property

bool: True if dtype_code encodes a string width (negative).

looks_sane property

Heuristic sanity check distinguishing a real channel record from leftover-garbage bytes.

Returns:

Type Description
bool

True if this record looks like real channel data rather than leftover-garbage bytes that happen to decode a clean printable name (observed for real in DB_Mag_833.gdb -- see docs/provenance/notes.md). Real records seen so far always have dtype either a known GS_* code (0-13) or a small negative string width, a format code in the known DB_CHAN_FORMAT_* range (0-6), and an array width of at least 1 (elements per fiducial).

Notes

The array-width test catches leftover records that pass the other two: 16 records in three real GSQ files (DB_AGG_1213, DB_Mag_1213, DB_Mag_1212) hold clean names -- fill-pattern names, projection-catalog names, second copies of real channel names -- with dtype 0 and array width 0 (or garbage), and own no data blob at all. Every genuine channel in the corpus has width 1 or more (docs/provenance/notes.md section 6.2d).

string_width property

int or None: Fixed on-disk width in bytes, for a string channel.

type_name property

str: Human-readable type name (a GS_* name, or "string[N]").

pygdb.LineRecord(index, offset, name, category_code, raw, name_is_clean=True) dataclass

One 128-byte line-table record.

Attributes:

Name Type Description
index int

0-based physical slot number -- this IS line_slot_index in the blob_index formula (BlobHeader.line_channel).

offset int

Byte offset of this record within the file.

name str

Line name.

category_code int or None

Raw int32 category code (see DB_CATEGORY_LINE_NAMES).

raw bytes

The record's raw, undecoded bytes.

name_is_clean bool, default True

See ChannelRecord.name_is_clean.

Notes

[LIKELY]/[UNKNOWN] -- much less firmly established than ChannelRecord: only the name (relative +32) and category code (relative +108) fields are decoded. The table is located exactly (exact_line_table_start, [CONFIRMED] on the corpus) and only falls back to the find_line_table heuristic scan when that cannot be validated. See docs/spec.md section 3.2 and docs/provenance/notes.md section 6.3.

eq=False keeps the default identity-based __eq__/__hash__ instead of dataclass's usual field-by-field one -- needed both to stay hashable (see ChannelRecord's docstring for why) and because GDB._calibrate_line_indices mutates .index in place on these after construction; a value-based __eq__/__hash__ pair would be actively wrong for an object whose fields change post-construction.

category_name property

str: Human-readable category name (a DB_CATEGORY_LINE_* name).

pygdb.BlobHeader(offset, n_pages, n_pages_dup, blob_index, timestamp, reserved_200, scale, row_count, gs_type_code) dataclass

The per-channel-per-line data block header.

Attributes:

Name Type Description
offset int

Absolute file offset of this header's first byte.

n_pages int

[CONFIRMED] -- this blob's total on-disk size, in pages (page_size from header_fields()).

n_pages_dup int

[LIKELY] -- always seen equal to n_pages.

blob_index int

[CONFIRMED] -- see line_channel.

timestamp int

[LIKELY], modern files only -- see Notes.

reserved_200 int

[UNKNOWN].

scale float

[LIKELY], modern files only.

row_count int

[CONFIRMED], modern files only (verified against real ground-truth-matching decoded values).

gs_type_code int

[CONFIRMED], modern files only (matches owning channel's own symbol-table dtype exactly).

Notes

[CONFIRMED] for fields up to and including blob_index (verified byte-exact on 5 real DB_COMP_NONE files via a whole-file, zero-error chain walk that lands exactly on each file's true size -- see docs/provenance/notes.md section 6.6). Fields from timestamp onward are only [LIKELY]/[UNKNOWN] and are known NOT to decode sensibly at these byte offsets in at least one real older (1991 GSQ) file -- kept here for the modern (2020 USGS) case where they were verified, not assumed general.

data_offset property

int: Absolute file offset where this blob's value data begins.

line_channel(chans_max)

Decompose blob_index into (line_slot_index, channel_slot_index).

Parameters:

Name Type Description Default
chans_max int

The file's channel-table capacity (from header_fields()).

required

Returns:

Type Description
tuple of (int, int)

(line_slot_index, channel_slot_index), both 0-based physical slot numbers in their respective symbol tables (same indexing as ChannelRecord.index and the line table walked ad hoc in docs/provenance/notes.md section 6.3), via the formula [CONFIRMED] in docs/provenance/notes.md section 6.6:

::

blob_index == line_slot_index * chans_max + channel_slot_index
Source code in pygdb/gdb_reader.py
def line_channel(self, chans_max: int):
    """
    Decompose `blob_index` into `(line_slot_index, channel_slot_index)`.

    Parameters
    ----------
    chans_max : int
        The file's channel-table capacity (from `header_fields()`).

    Returns
    -------
    tuple of (int, int)
        `(line_slot_index, channel_slot_index)`, both 0-based
        physical slot numbers in their respective symbol tables
        (same indexing as `ChannelRecord.index` and the line table
        walked ad hoc in docs/provenance/notes.md section 6.3),
        via the formula **[CONFIRMED]** in
        docs/provenance/notes.md section 6.6:

        ::

            blob_index == line_slot_index * chans_max + channel_slot_index
    """
    return divmod(self.blob_index, chans_max)

pygdb.BlobDirectory(data_slots, entries) dataclass

The persisted blob directory: which blob is the live one for each slot.

Attributes:

Name Type Description
data_slots int

Number of directory slots that address (line, channel) data blobs (header word 48, lines_max * chans_max); slot i is the blob whose blob_index is i.

entries dict of {int : (int, int)}

Every non-zero data slot as (32-bit word, n_pages), keyed by slot. A zero slot is simply absent from this dict.

Notes

[CONFIRMED] layout, [LIKELY] interpretation (docs/spec.md section 2.2; docs/provenance/notes.md section 6.1c). The directory is an array of 6-byte slots starting at file offset 280. A live entry is (0x80000000 | start page, n_pages) where the start page is relative to the first blob ((offset - blob region start) / page_size), optionally with the 0x40000000 bit, which flips each time the blob is rewritten. (Every data entry in the first 22 corpus files read 0x8; the first data entry with 0xC, in USGS OFR 2011-1270 Kalay_nk.gdb, is the live copy -- the other copy is on the free list.) Across the real corpus it addressed 100% of the real (line, channel) blobs of 20 of 22 files, and for every duplicated pair with an independent oracle (13 of 13) it pointed at the correct copy. A slot that is all-zero belongs to a blob the file does not list as live.

resolve(blob_index, blobs_by_offset, first_offset, page_size)

Look up the live blob for blob_index.

Parameters:

Name Type Description Default
blob_index int

The slot, line_slot * chans_max + channel_slot.

required
blobs_by_offset dict of {int : BlobHeader}

Every blob of the chain walk, keyed by absolute offset.

required
first_offset int

Absolute offset of the first blob (blob_region_start).

required
page_size int

The file's page size.

required

Returns:

Name Type Description
status str

"live": the entry's start page lands on a walked blob header whose blob_index equals blob_index and whose n_pages equals the entry's page count (the strict check). "absent": the slot is all-zero. "invalid": the slot is non-zero but fails the strict check. "outside": blob_index is not a data slot at all (an administrative slot), so the directory says nothing about it.

blob BlobHeader or None

The live blob for "live", otherwise None.

Source code in pygdb/gdb_reader.py
def resolve(self, blob_index: int, blobs_by_offset: dict, first_offset: int, page_size: int):
    """
    Look up the live blob for `blob_index`.

    Parameters
    ----------
    blob_index : int
        The slot, `line_slot * chans_max + channel_slot`.
    blobs_by_offset : dict of {int : BlobHeader}
        Every blob of the chain walk, keyed by absolute `offset`.
    first_offset : int
        Absolute offset of the first blob (`blob_region_start`).
    page_size : int
        The file's page size.

    Returns
    -------
    status : str
        `"live"`: the entry's start page lands on a walked blob
        header whose `blob_index` equals `blob_index` and whose
        `n_pages` equals the entry's page count (the strict
        check). `"absent"`: the slot is all-zero. `"invalid"`: the
        slot is non-zero but fails the strict check. `"outside"`:
        `blob_index` is not a data slot at all (an administrative
        slot), so the directory says nothing about it.
    blob : BlobHeader or None
        The live blob for `"live"`, otherwise `None`.
    """
    if not 0 <= blob_index < self.data_slots:
        return DIRECTORY_OUTSIDE, None
    entry = self.entries.get(blob_index)
    if entry is None:
        return DIRECTORY_ABSENT, None
    word, n_pages = entry
    if not word & _DIRECTORY_LIVE_BIT:
        return DIRECTORY_INVALID, None
    blob = blobs_by_offset.get(first_offset + (word & _DIRECTORY_PAGE_MASK) * page_size)
    if blob is None or blob.blob_index != blob_index or blob.n_pages != n_pages:
        return DIRECTORY_INVALID, None
    return DIRECTORY_LIVE, blob

.gdb reader functions

pygdb.check_magic(data)

Check whether data starts with the format's magic bytes.

Parameters:

Name Type Description Default
data bytes

The file's leading bytes (at least 4).

required

Returns:

Type Description
bool

True if data starts with the 4-byte "!CBD" magic.

Notes

[CONFIRMED] against 9/9 real files across two independent sources (2020 USGS Mojave survey, 1990s-2020s GSQ Queensland surveys from three different TEM systems/vendors).

The full 16-byte HEADER_SIGNATURE is only [LIKELY] -- it matched exactly in 8/9 real files, but one real file (DB_Mag_Elaine_1003.gdb, from GSQ's Mount Gordon delivery) has f0 f0 f0 f0 at bytes 4-7 instead of the usual 00 00 00 00. That file is otherwise structurally normal (chans_max/users_max/page_size all decode sanely), so this looks like a real, if rare, variation in that sub-block rather than a different format entirely -- flagged [UNKNOWN] in docs/provenance/notes.md. Only the 4-byte magic is treated as a hard requirement here; the rest of the signature is reported separately (magic_signature_matches_common_case).

Source code in pygdb/gdb_reader.py
def check_magic(data: bytes) -> bool:
    """
    Check whether `data` starts with the format's magic bytes.

    Parameters
    ----------
    data : bytes
        The file's leading bytes (at least 4).

    Returns
    -------
    bool
        True if `data` starts with the 4-byte `"!CBD"` magic.

    Notes
    -----
    **[CONFIRMED]** against 9/9 real files across two independent
    sources (2020 USGS Mojave survey, 1990s-2020s GSQ Queensland
    surveys from three different TEM systems/vendors).

    The full 16-byte `HEADER_SIGNATURE` is only **[LIKELY]** -- it
    matched exactly in 8/9 real files, but one real file
    (DB_Mag_Elaine_1003.gdb, from GSQ's Mount Gordon delivery) has
    `f0 f0 f0 f0` at bytes 4-7 instead of the usual `00 00 00 00`.
    That file is otherwise structurally normal
    (chans_max/users_max/page_size all decode sanely), so this looks
    like a real, if rare, variation in that sub-block rather than a
    different format entirely -- flagged **[UNKNOWN]** in
    docs/provenance/notes.md. Only the 4-byte magic is treated as a
    hard requirement here; the rest of the signature is reported
    separately (`magic_signature_matches_common_case`).
    """
    return data[:4] == MAGIC

pygdb.header_fields(data)

Extract the header int32 fields whose approximate meaning is known.

Parameters:

Name Type Description Default
data bytes

The file's header bytes.

required

Returns:

Type Description
dict

Maps each of "chans_max" (word 24), "blobs_max" (28), "lines_max" (36), "users_max" (40), "index_slots" (44, the total number of blob-directory slots), "data_slots" (48, the number of those that address (line, channel) data blobs), "page_size" (100), and "comp_level" (120) to its decoded int32 value. See docs/spec.md section 2 for the full table with a confidence rating per word, and docs/provenance/notes.md section 6.⅙.1b for the derivations and the still-unknown offsets.

Warns:

Type Description
GDBParseWarning

If data is truncated: any field that can't be read (not enough bytes at its offset) is set to None in the returned dict rather than raising. Callers that need a field should check for None before using it (every function in this module that consumes header_fields() output does).

Notes

comp_level==1 (DB_COMP_SPEED) does NOT mean the payload is zlib -- confirmed it is NOT (docs/provenance/notes.md section 6.5b), it's canonical LZRW1 (section 6.5c). comp_level==2 (DB_COMP_SIZE) IS confirmed real zlib.

Source code in pygdb/gdb_reader.py
def header_fields(data: bytes) -> dict:
    """
    Extract the header int32 fields whose approximate meaning is known.

    Parameters
    ----------
    data : bytes
        The file's header bytes.

    Returns
    -------
    dict
        Maps each of `"chans_max"` (word 24), `"blobs_max"` (28),
        `"lines_max"` (36), `"users_max"` (40), `"index_slots"` (44,
        the total number of blob-directory slots), `"data_slots"` (48,
        the number of those that address (line, channel) data blobs),
        `"page_size"` (100), and `"comp_level"` (120) to its decoded
        int32 value. See docs/spec.md section 2 for the full table
        with a confidence rating per word, and
        docs/provenance/notes.md section 6.1/6.1b for the derivations
        and the still-unknown offsets.

    Warns
    -----
    GDBParseWarning
        If `data` is truncated: any field that can't be read (not
        enough bytes at its offset) is set to `None` in the returned
        dict rather than raising. Callers that need a field should
        check for `None` before using it (every function in this
        module that consumes `header_fields()` output does).

    Notes
    -----
    `comp_level==1` (`DB_COMP_SPEED`) does NOT mean the payload is
    zlib -- confirmed it is NOT (docs/provenance/notes.md section
    6.5b), it's canonical LZRW1 (section 6.5c). `comp_level==2`
    (`DB_COMP_SIZE`) IS confirmed real zlib.
    """
    result = {}
    for name, offset in (("chans_max", 24), ("blobs_max", 28), ("lines_max", 36),
                          ("users_max", 40), ("index_slots", 44), ("data_slots", 48),
                          ("page_size", 100), ("comp_level", 120)):
        try:
            result[name] = struct.unpack_from("<i", data, offset)[0]
        except struct.error:
            _warn(
                f"header truncated: only {len(data)} byte(s) available, not enough "
                f"to read '{name}' at offset {offset} -- returning None for it"
            )
            result[name] = None
    # comp_level==1 (DB_COMP_SPEED) does NOT mean the payload is zlib --
    # confirmed it is NOT (docs/provenance/notes.md section 6.5b), it's canonical LZRW1
    # (section 6.5c). comp_level==2 (DB_COMP_SIZE) IS confirmed real zlib.
    return result

pygdb.find_channel_table(data, search_window=(0, None))

Locate the start of the channel symbol table.

Parameters:

Name Type Description Default
data bytes

Buffer to search (typically the start of the file).

required
search_window tuple of (int, int or None)

(lo, hi) byte range within data to search; hi=None means the end of data.

(0, None)

Returns:

Type Description
int

Byte offset of the channel table's first record.

Raises:

Type Description
ValueError

If no "SUPER"/"super" occurrence in the search window implies a valid channel table.

Notes

Strategy [CONFIRMED against 2 real 2020 USGS files, RE-CONFIRMED -- with one revision -- against 5 more real 1990s-2020s GSQ files, see docs/provenance/notes.md/docs/provenance/log.md "pressure test" round]: search for the default super-user name (from GXDB.create()'s documented default super="SUPER"). The channel table is found to occupy exactly chans_max consecutive 128-byte records immediately before the user table, i.e.

::

channel_table_start == offset_of(super_name) - 8 - chans_max*128

This isn't a generic file-format constant we can hardcode a single offset for -- it depends on chans_max, which itself varies between files -- so we compute it.

REVISION from the original derivation: the two 2020 USGS files both had the default super-user name stored as literal uppercase ASCII "SUPER". Five real 1990s-2020s GSQ files instead have it stored as lowercase "super" -- confirmed to be the same structural pattern (same 128-byte-per-record math, same position relative to the channel table) once the case is corrected, not a different layout. Search for both cases. (One of the GSQ files, DB_Mag_833.gdb, also demonstrated that the literal string can coincidentally appear elsewhere in a file, e.g. inside embedded metadata blobs, and that the word "super"/"SUPER" appearing is not on its own sufficient -- a naive first-match there pointed at a bogus offset. Confirmed correct instead via an independent generic 128-byte-periodicity scan that landed on the identical answer once cross-checked.)

Every occurrence of "SUPER"/"super" is tried and the first one whose implied table start decodes a clean (NUL-terminated, printable) channel name is used.

Source code in pygdb/gdb_reader.py
def find_channel_table(data: bytes, search_window=(0, None)) -> int:
    """
    Locate the start of the channel symbol table.

    Parameters
    ----------
    data : bytes
        Buffer to search (typically the start of the file).
    search_window : tuple of (int, int or None), default (0, None)
        `(lo, hi)` byte range within `data` to search; `hi=None`
        means the end of `data`.

    Returns
    -------
    int
        Byte offset of the channel table's first record.

    Raises
    ------
    ValueError
        If no `"SUPER"`/`"super"` occurrence in the search window
        implies a valid channel table.

    Notes
    -----
    Strategy **[CONFIRMED against 2 real 2020 USGS files, RE-CONFIRMED
    -- with one revision -- against 5 more real 1990s-2020s GSQ files,
    see docs/provenance/notes.md/docs/provenance/log.md "pressure
    test" round]**: search for the default super-user name (from
    `GXDB.create()`'s documented default `super="SUPER"`). The channel
    table is found to occupy exactly `chans_max` consecutive 128-byte
    records immediately before the user table, i.e.

    ::

        channel_table_start == offset_of(super_name) - 8 - chans_max*128

    This isn't a generic file-format constant we can hardcode a single
    offset for -- it depends on `chans_max`, which itself varies
    between files -- so we compute it.

    REVISION from the original derivation: the two 2020 USGS files
    both had the default super-user name stored as literal uppercase
    ASCII "SUPER". Five real 1990s-2020s GSQ files instead have it
    stored as lowercase "super" -- confirmed to be the *same*
    structural pattern (same 128-byte-per-record math, same position
    relative to the channel table) once the case is corrected, not a
    different layout. Search for both cases. (One of the GSQ files,
    DB_Mag_833.gdb, also demonstrated that the literal string can
    coincidentally appear elsewhere in a file, e.g. inside embedded
    metadata blobs, and that the *word* "super"/"SUPER" appearing is
    not on its own sufficient -- a naive first-match there pointed at
    a bogus offset. Confirmed correct instead via an independent
    generic 128-byte-periodicity scan that landed on the identical
    answer once cross-checked.)

    Every occurrence of "SUPER"/"super" is tried and the first one
    whose implied table start decodes a *clean* (NUL-terminated,
    printable) channel name is used.
    """
    lo, hi = search_window
    if hi is None:
        hi = len(data)
    chans_max = header_fields(data)["chans_max"]

    start = lo
    while True:
        idx_upper = data.find(b"SUPER", start, hi)
        idx_lower = data.find(b"super", start, hi)
        candidates = [i for i in (idx_upper, idx_lower) if i != -1]
        super_idx = min(candidates) if candidates else -1
        if super_idx == -1:
            raise ValueError(
                "could not find a 'SUPER' user record that implies a valid "
                "channel table in the search window; try widening `search_window`"
            )
        super_rec_start = super_idx - 8
        table_start = super_rec_start - chans_max * SYMBOL_RECORD_SIZE
        if table_start >= 0:
            raw = data[table_start : table_start + SYMBOL_RECORD_SIZE]
            if len(raw) == SYMBOL_RECORD_SIZE:
                name, is_clean = _read_name(raw, 8)
                if is_clean and name:
                    return table_start
        start = super_idx + 1

pygdb.read_channels(path)

Decode the channel symbol table.

Parameters:

Name Type Description Default
path str

Path to the .gdb file.

required

Returns:

Type Description
list of ChannelRecord

Every channel successfully decoded, in table order.

Warns:

Type Description
GDBParseWarning

A file that isn't a real .gdb (bad magic), has a truncated header, has no locatable channel table, or has a channel table that's cut off partway through all result in this warning plus whatever channels were successfully decoded before the problem (an empty list in the first three cases, since nothing was decodable yet; a real, non-empty, shorter-than-chans_max list in the last case). Never raises for these -- see GDBParseWarning's docstring.

Source code in pygdb/gdb_reader.py
def read_channels(path: str) -> List[ChannelRecord]:
    """
    Decode the channel symbol table.

    Parameters
    ----------
    path : str
        Path to the `.gdb` file.

    Returns
    -------
    list of ChannelRecord
        Every channel successfully decoded, in table order.

    Warns
    -----
    GDBParseWarning
        A file that isn't a real `.gdb` (bad magic), has a truncated
        header, has no locatable channel table, or has a channel
        table that's cut off partway through all result in this
        warning plus whatever channels *were* successfully decoded
        before the problem (an empty list in the first three cases,
        since nothing was decodable yet; a real, non-empty,
        shorter-than-`chans_max` list in the last case). Never raises
        for these -- see `GDBParseWarning`'s docstring.
    """
    with open(path, "rb") as f:
        # Reading the whole file is wasteful for a 700MB+ real survey
        # database, but the symbol table's exact byte extent isn't fully
        # pinned down yet (docs/provenance/notes.md 6.1), so for correctness this reads
        # generously. A production version should read a memory-mapped
        # view instead -- left as a TODO once the header's table-size
        # field (offset 104, currently [UNKNOWN]) is confirmed.
        header = f.read(4096)
        if not check_magic(header):
            _warn(f"{path}: does not start with the expected '!CBD' magic -- "
                  f"not a recognized .gdb file, returning no channels")
            return []
        fields = header_fields(header)
        if fields["chans_max"] is None:
            _warn(f"{path}: header too short to read chans_max -- returning no channels")
            return []

        f.seek(0, 2)
        size = f.tell()
        f.seek(0)
        # Only need enough of the file to reach the channel + user tables.
        # Observed table offsets range from ~130KB to ~580KB across 9 real
        # files so far, but read generously (all of a file up to 200MB,
        # else the first 20MB) since the exact extent isn't pinned down.
        data = f.read(size if size <= 200_000_000 else 20_000_000)

    try:
        table_start = find_channel_table(data)
    except ValueError as e:
        _warn(f"{path}: could not locate the channel symbol table ({e}) -- "
              f"returning no channels")
        return []
    chans_max = fields["chans_max"]

    channels = []
    for i in range(chans_max):
        rec_start = table_start + i * SYMBOL_RECORD_SIZE
        if rec_start + SYMBOL_RECORD_SIZE > len(data):
            _warn(
                f"{path}: channel table truncated at record {i} of {chans_max} "
                f"(need bytes up to {rec_start + SYMBOL_RECORD_SIZE}, only "
                f"{len(data)} were read/available) -- returning the "
                f"{len(channels)} channel(s) decoded so far"
            )
            break
        rec = _parse_channel_record(data, rec_start, i)
        if not rec.name:
            continue  # cleanly empty/unused slot (NUL name, zeroed record)
        if not rec.name_is_clean:
            # Unused capacity slot with leftover/uninitialized bytes rather
            # than a clean NUL name -- observed in at least one real 1990s
            # file (see _read_name docstring / docs/provenance/notes.md). Not a real channel.
            continue
        if not rec.looks_sane:
            # A NUL-terminated printable "name" can still show up by pure
            # coincidence inside leftover garbage bytes in an unused slot
            # (observed in DB_Mag_833.gdb: "L2161" and "1", both leftover
            # fragments of an embedded projection-name blob that happened
            # to land in unused channel-table capacity). Real channel
            # records always have a dtype matching a known GS_* code or a
            # small negative string width, and a format code in the known
            # DB_CHAN_FORMAT_* range -- garbage doesn't. See docs/provenance/notes.md.
            continue
        channels.append(rec)
    unseen.check_channels(path, channels)
    if fields["users_max"] is not None:
        # The user table directly follows the channel table (docs/spec.md section 2.1).
        unseen.check_users(path, data, table_start + chans_max * SYMBOL_RECORD_SIZE, fields["users_max"])
    return channels

pygdb.exact_line_table_start(data, lines_max)

Compute the line table's start from the header and the channel table.

Parameters:

Name Type Description Default
data bytes

The file's bytes up to at least the end of the channel table (read_lines passes everything before the blob region).

required
lines_max int or None

The line-table capacity (header word 36); None or non-positive means unknown.

required

Returns:

Type Description
int or None

Byte offset of the line table's first record, or None if the arithmetic cannot be validated (unknown lines_max, no locatable channel table, a start before the header ends, or no line-shaped record anywhere in the computed table) -- the caller then falls back to find_line_table.

Notes

The line table is lines_max 128-byte slots followed by a 24-byte gap and then the channel table, so its start is find_channel_table(data) - 24 - lines_max * 128. [CONFIRMED] on all 23 real files examined (docs/spec.md section 2.1; docs/provenance/notes.md section 6.1b): the computed start is always a slot boundary that matches the heuristic's start, or is an earlier slot the heuristic missed. The at-least-one-line-record check only guards a file whose layout does not follow this arithmetic; it is not needed for any real file seen.

Source code in pygdb/gdb_reader.py
def exact_line_table_start(data: bytes, lines_max: Optional[int]) -> Optional[int]:
    """
    Compute the line table's start from the header and the channel table.

    Parameters
    ----------
    data : bytes
        The file's bytes up to at least the end of the channel table
        (`read_lines` passes everything before the blob region).
    lines_max : int or None
        The line-table capacity (header word 36); `None` or non-positive
        means unknown.

    Returns
    -------
    int or None
        Byte offset of the line table's first record, or `None` if the
        arithmetic cannot be validated (unknown `lines_max`, no
        locatable channel table, a start before the header ends, or no
        line-shaped record anywhere in the computed table) -- the
        caller then falls back to `find_line_table`.

    Notes
    -----
    The line table is `lines_max` 128-byte slots followed by a 24-byte
    gap and then the channel table, so its start is
    `find_channel_table(data) - 24 - lines_max * 128`. **[CONFIRMED]**
    on all 23 real files examined (docs/spec.md section 2.1;
    docs/provenance/notes.md section 6.1b): the computed start is
    always a slot boundary that matches the heuristic's start, or is an
    earlier slot the heuristic missed. The at-least-one-line-record
    check only guards a file whose layout does not follow this
    arithmetic; it is not needed for any real file seen.
    """
    if not lines_max or lines_max <= 0:
        return None
    try:
        channel_start = find_channel_table(data)
    except ValueError:
        return None
    start = channel_start - _LINE_TABLE_GAP - lines_max * SYMBOL_RECORD_SIZE
    if start < 256:
        return None
    for i in range(lines_max):
        rec = _parse_line_record(data, start + i * SYMBOL_RECORD_SIZE, i)
        if rec.name and rec.name_is_clean and rec.category_code in DB_CATEGORY_LINE_NAMES:
            return start
    return None

pygdb.find_line_table(data, search_window=(128, None))

Locate the start of the line symbol table.

Parameters:

Name Type Description Default
data bytes

Buffer to search.

required
search_window tuple of (int, int or None)

(lo, hi) byte range within data to search; hi=None means the end of data. Callers should generally narrow hi to blob_region_start(data) when known, since every real file examined has its symbol tables (line, channel, user) entirely before the blob region, and narrowing avoids false-positive matches inside actual channel data.

(128, None)

Returns:

Type Description
int

Byte offset of the line table's first record.

Raises:

Type Description
ValueError

If no line-record-shaped data (clean name + a known category code) is found in the search window.

Notes

Prefer exact_line_table_start (via read_lines): the line table's start follows exactly from lines_max and the channel table's position (docs/spec.md section 2.1), which this heuristic predates. This scan is now only the fallback for a file where that arithmetic cannot be validated.

Unlike find_channel_table, there's no known default-name anchor (the line table has nothing analogous to the channel table's "SUPER" user record immediately after it), so this is a heuristic [LIKELY] scan, not a structurally-proven technique: it looks for a run of 128-byte records whose relative +32 field looks like a clean, NUL-terminated, printable line name and whose relative +108 category field matches one of the two confirmed real values (100=NORMAL/FLIGHT, 200=GROUP), then returns the earliest such record in the run with the most hits at a consistent 128-byte phase.

Known limitation, found by real-file testing (fixed by using exact_line_table_start instead, not here): if a table's true first slot(s) don't carry a category code in {100, 200}, this returns a start that's one or more slots too late -- every subsequent LineRecord.index is then off by that same fixed amount, which breaks blob_index lookups by line name. Observed for real on a GSQ file (rm001141): physical slot 0 is a genuine, named record ("L0") with category 65636 ([GUESS]: 65536 + 100, plausibly "a NORMAL line that was since cleared," not confirmed), which this function doesn't recognize, so it starts the table one slot late. A generic backward-scan fix was tried and rejected: "keep walking backward while the name field still looks clean" massively over-extends on at least one real file (walked 30+ slots into what turned out to be unrelated, legitimately-empty space before the real table). GDB (in gdb.py) instead cross-validates and corrects this against the actual blob chain, which is a strictly stronger signal than anything available from the symbol-table bytes alone -- prefer it over calling this function directly when correct line-indexed data access matters, not just names.

Source code in pygdb/gdb_reader.py
def find_line_table(data: bytes, search_window: Tuple[int, Optional[int]] = (128, None)) -> int:
    """
    Locate the start of the line symbol table.

    Parameters
    ----------
    data : bytes
        Buffer to search.
    search_window : tuple of (int, int or None), default (128, None)
        `(lo, hi)` byte range within `data` to search; `hi=None`
        means the end of `data`. Callers should generally narrow `hi`
        to `blob_region_start(data)` when known, since every real
        file examined has its symbol tables (line, channel, user)
        entirely before the blob region, and narrowing avoids
        false-positive matches inside actual channel data.

    Returns
    -------
    int
        Byte offset of the line table's first record.

    Raises
    ------
    ValueError
        If no line-record-shaped data (clean name + a known category
        code) is found in the search window.

    Notes
    -----
    **Prefer `exact_line_table_start`** (via `read_lines`): the line
    table's start follows exactly from `lines_max` and the channel
    table's position (docs/spec.md section 2.1), which this heuristic
    predates. This scan is now only the fallback for a file where that
    arithmetic cannot be validated.

    Unlike `find_channel_table`, there's no known default-name anchor
    (the line table has nothing analogous to the channel table's
    "SUPER" user record immediately after it), so this is a
    heuristic **[LIKELY]** scan, not a structurally-proven technique:
    it looks for a run of 128-byte records
    whose relative +32 field looks like a clean, NUL-terminated,
    printable line name and whose relative +108 category field
    matches one of the two confirmed real values (100=NORMAL/FLIGHT,
    200=GROUP), then returns the earliest such record in the run with
    the most hits at a consistent 128-byte phase.

    **Known limitation, found by real-file testing (fixed by using
    `exact_line_table_start` instead, not here):** if a table's true first slot(s) don't carry a category
    code in {100, 200}, this returns a start that's one or more slots
    too late -- every subsequent `LineRecord.index` is then off by
    that same fixed amount, which breaks blob_index lookups by line
    name. Observed for real on a GSQ file (`rm001141`): physical slot
    0 is a genuine, named record (`"L0"`) with category `65636`
    (**[GUESS]**: `65536 + 100`, plausibly "a NORMAL line that was
    since cleared," not confirmed), which this function doesn't
    recognize, so it starts the table one slot late. A generic
    backward-scan fix was tried and rejected: "keep walking backward
    while the name field still looks clean" massively over-extends on
    at least one real file (walked 30+ slots into what turned out to
    be unrelated, legitimately-empty space before the real table).
    `GDB` (in `gdb.py`) instead cross-validates and corrects this
    against the actual blob chain, which is a strictly stronger
    signal than anything available from the symbol-table bytes alone
    -- prefer it over calling this function directly when correct
    line-indexed data access matters, not just names.
    """
    lo, hi = search_window
    if hi is None:
        hi = len(data)

    phase_hits = {}
    for m in _NAME_LIKE_RE.finditer(data, lo, hi):
        rec_start = m.start() - 32
        if rec_start < lo or rec_start + SYMBOL_RECORD_SIZE > hi:
            continue
        raw = data[rec_start : rec_start + SYMBOL_RECORD_SIZE]
        try:
            category_code = struct.unpack_from("<i", raw, 108)[0]
        except struct.error:
            continue
        if category_code not in DB_CATEGORY_LINE_NAMES:
            continue
        phase_hits.setdefault(rec_start % SYMBOL_RECORD_SIZE, []).append(rec_start)

    if not phase_hits:
        raise ValueError(
            "no line-record-shaped data (clean name + a known category "
            "code) found in the search window"
        )
    best_phase = max(phase_hits, key=lambda p: len(phase_hits[p]))
    return min(phase_hits[best_phase])

pygdb.read_lines(path)

Decode the line symbol table.

Parameters:

Name Type Description Default
path str

Path to the .gdb file.

required

Returns:

Type Description
list of LineRecord

Every line successfully decoded, in table order.

Warns:

Type Description
GDBParseWarning

A bad magic, truncated header, or unlocatable line table all return [] with this warning rather than raising.

Notes

The table is located exactly whenever possible (see exact_line_table_start): it holds lines_max (header word 36) 128-byte slots and ends 24 bytes before the channel table, so its start is channel_table_start - 24 - lines_max * 128 ([CONFIRMED] on every real file in the corpus, docs/spec.md section 2.1). Every one of the lines_max slots is then examined, and LineRecord.index is the true slot number, so LineRecord.index is exact and blob_index lookups keyed on it are right.

Only when that arithmetic cannot be validated (a header without lines_max, or no line-shaped record at the computed position) does this fall back to the older heuristic (see find_line_table): [LIKELY], less firmly established, reading forward until 8 consecutive records fail to look like either a populated line record or clean unused capacity. In that fallback LineRecord.index can be off by a small, fixed amount when the scan starts one or more slots late -- .name is still correct but .index is wrong. GDB (in gdb.py) corrects this against the actual blob chain only in that case.

A populated slot whose category is neither 100 nor 200 (for example the 65636 slot 0 of one 1991 file) is not returned, as before -- but, unlike the fallback, it no longer shifts the indices of the lines after it.

Source code in pygdb/gdb_reader.py
def read_lines(path: str) -> List[LineRecord]:
    """
    Decode the line symbol table.

    Parameters
    ----------
    path : str
        Path to the `.gdb` file.

    Returns
    -------
    list of LineRecord
        Every line successfully decoded, in table order.

    Warns
    -----
    GDBParseWarning
        A bad magic, truncated header, or unlocatable line table all
        return `[]` with this warning rather than raising.

    Notes
    -----
    The table is located **exactly** whenever possible (see
    `exact_line_table_start`): it holds `lines_max` (header word 36)
    128-byte slots and ends 24 bytes before the channel table, so its
    start is `channel_table_start - 24 - lines_max * 128`
    (**[CONFIRMED]** on every real file in the corpus, docs/spec.md
    section 2.1). Every one of the `lines_max` slots is then examined,
    and `LineRecord.index` is the true slot number, so
    `LineRecord.index` is exact and blob_index lookups keyed on it
    are right.

    Only when that arithmetic cannot be validated (a header without
    `lines_max`, or no line-shaped record at the computed position) does
    this fall back to the older heuristic (see `find_line_table`):
    **[LIKELY]**, less firmly established, reading forward until 8
    consecutive records fail to look like either a populated line
    record or clean unused capacity. In that fallback
    **`LineRecord.index` can be off by a small, fixed amount** when the
    scan starts one or more slots late -- `.name` is still correct but
    `.index` is wrong. `GDB` (in `gdb.py`) corrects this against the
    actual blob chain only in that case.

    A populated slot whose category is neither `100` nor `200` (for
    example the `65636` slot 0 of one 1991 file) is not returned, as
    before -- but, unlike the fallback, it no longer shifts the
    indices of the lines after it.
    """
    return _read_lines(path)[0]

pygdb.iter_blobs(path, max_blobs=None)

Walk the self-describing blob chain.

Walks from the start of the blob region to end of file (or max_blobs, or a framing anomaly it cannot step past). A page that does not start a blob is skipped: the walk resumes at the next page that does.

Parameters:

Name Type Description Default
path str

Path to the .gdb file.

required
max_blobs int

Stop after yielding this many blobs, if given.

None

Yields:

Type Description
BlobHeader

Each blob header, in on-disk order.

Warns:

Type Description
GDBParseWarning

Reaching the file's true end cleanly is silent (the expected, common case). Pages skipped because they do not start a blob get one summary warning (count and offsets). A magic mismatch with no later blob page, a non-positive n_pages, a file that ends mid-header, or landing short of true EOF by less than one full header (i.e. real leftover bytes, not enough to be read at all) are all real anomalies and each gets its own specific warning identifying the offset and how many blobs were walked first -- so a caller can tell "the chain looked completely normal and just ended" from "something didn't fit the confirmed structure" without having to guess from the return value alone. Never raises for a bad/truncated file; only for a real precondition problem (can't even open path, propagated normally from open).

Notes

[CONFIRMED] end-to-end (zero framing errors, landing exactly on the true file size) on 20 real files spanning all 3 agencies this project has files from and all three DB_COMP_* compression modes, 2MB to 1.93GB -- see docs/provenance/notes.md section 6.6b/6.6d/6.9. One later real file (OpenEI BRIDGE GP_Master_Gravity_11082023.gdb, 2023) has pages between blobs that are not blobs: one page of leftover data at the region start and a run of 96 all-zero pages. Its blobs are still contiguous around those gaps, and the resynchronizing walk lands exactly on the file's end.

As a generator, this already "returns partial results" in the most natural way possible: whatever's been yielded before a problem is hit stays with the caller (a for blob in iter_blobs(path): ... loop simply ends, keeping everything already processed) -- nothing is lost by stopping early.

Source code in pygdb/gdb_reader.py
def iter_blobs(path: str, max_blobs: Optional[int] = None):
    """
    Walk the self-describing blob chain.

    Walks from the start of the blob region to end of file (or
    `max_blobs`, or a framing anomaly it cannot step past). A page that
    does not start a blob is skipped: the walk resumes at the next page
    that does.

    Parameters
    ----------
    path : str
        Path to the `.gdb` file.
    max_blobs : int, optional
        Stop after yielding this many blobs, if given.

    Yields
    ------
    BlobHeader
        Each blob header, in on-disk order.

    Warns
    -----
    GDBParseWarning
        Reaching the file's true end cleanly is silent (the expected,
        common case). Pages skipped because they do not start a blob get
        one summary warning (count and offsets). A magic mismatch with no
        later blob page, a non-positive `n_pages`,
        a file that ends mid-header, or landing short of true EOF by
        less than one full header (i.e. real leftover bytes, not
        enough to be read at all) are all real anomalies and each gets
        its own specific warning identifying the offset and how many
        blobs were walked first -- so a caller can tell "the chain
        looked completely normal and just ended" from "something
        didn't fit the confirmed structure" without having to guess
        from the return value alone. Never raises for a bad/truncated
        file; only for a real precondition problem (can't even open
        `path`, propagated normally from `open`).

    Notes
    -----
    **[CONFIRMED]** end-to-end (zero framing errors, landing exactly
    on the true file size) on 20 real files spanning all 3 agencies
    this project has files from and all three `DB_COMP_*` compression
    modes, 2MB to 1.93GB -- see docs/provenance/notes.md section
    6.6b/6.6d/6.9. One later real file (OpenEI BRIDGE
    `GP_Master_Gravity_11082023.gdb`, 2023) has pages between blobs
    that are not blobs: one page of leftover data at the region start
    and a run of 96 all-zero pages. Its blobs are still contiguous
    around those gaps, and the resynchronizing walk lands exactly on
    the file's end.

    As a generator, this already "returns partial results" in the most
    natural way possible: whatever's been yielded before a problem is
    hit stays with the caller (a `for blob in iter_blobs(path): ...`
    loop simply ends, keeping everything already processed) -- nothing
    is lost by stopping early.
    """
    with open(path, "rb") as f:
        header = f.read(128)
        if not check_magic(header):
            _warn(f"{path}: does not start with the expected '!CBD' magic -- "
                  f"no blobs to walk")
            return
        off = blob_region_start(header)
        if off is None:
            _warn(f"{path}: could not determine the blob region start -- "
                  f"no blobs to walk")
            return
        page_size = struct.unpack_from("<i", header, 100)[0]
        f.seek(0, 2)
        size = f.tell()
        if off > size:
            _warn(
                f"{path}: computed blob region start ({off}) is past the end "
                f"of the file ({size} byte(s)) -- file is likely severely "
                f"truncated; no blobs to walk"
            )
            return
        f.seek(off)
        n = 0
        skipped_runs = []  # [(first offset, page count)] of pages that are not blobs
        while off + BLOB_HEADER_SIZE <= size:
            if max_blobs is not None and n >= max_blobs:
                return
            raw = f.read(BLOB_HEADER_SIZE)
            blob = _parse_blob_header(raw, off)
            if blob is None:
                if len(raw) < BLOB_HEADER_SIZE:
                    _warn(
                        f"{path}: blob chain ends mid-header at offset {off} "
                        f"(only {len(raw)} of {BLOB_HEADER_SIZE} expected "
                        f"byte(s) available) after {n} blob(s) successfully "
                        f"walked -- file is likely truncated; returning the "
                        f"{n} blob(s) already yielded"
                    )
                    return
                # A page that is not a blob header. Real files can hold
                # such pages between blobs -- leftover data or never-written
                # zero pages (docs/spec.md section 6.2) -- so resynchronize on
                # the next page that starts with the blob magic.
                resume = _next_blob_page(f, off + page_size, page_size, size)
                if resume is None:
                    _warn(
                        f"{path}: blob magic mismatch at offset {off} "
                        f"(got {raw[:4].hex()}, expected {BLOB_MAGIC.hex()}) "
                        f"after {n} blob(s) successfully walked, and no later "
                        f"page starts a blob -- stopping the chain walk here "
                        f"and returning the {n} blob(s) already yielded"
                    )
                    return
                skipped_runs.append((off, (resume - off) // page_size))
                off = resume
                f.seek(off)
                continue
            if blob.n_pages <= 0:
                _warn(
                    f"{path}: blob at offset {off} (blob_index={blob.blob_index}) "
                    f"has a non-positive n_pages ({blob.n_pages}) after {n} "
                    f"blob(s) successfully walked -- cannot safely continue "
                    f"(don't know how far to skip to find the next header); "
                    f"returning the {n} blob(s) already yielded"
                )
                return
            # NOTE: n_pages_dup (relative +8) is NOT always equal to n_pages
            # (relative +4) -- confirmed on real Ontario GDS1251 files
            # (MLGRAV.gdb/MLMAG.gdb), where a small number of "reserved/
            # administrative" blobs (same class flagged [UNKNOWN] elsewhere
            # in this section -- out-of-range line index, gs_type_code
            # reading the same 4670802 constant) have n_pages_dup != n_pages.
            # Directly verified: n_pages (not n_pages_dup) is the field that
            # correctly lands on the next real blob header every time -- an
            # earlier version of this function required the two to match and
            # broke immediately on these files as a result. Trust n_pages
            # alone; n_pages_dup is kept on BlobHeader for whoever wants to
            # investigate what it actually means.
            yield blob
            skip = blob.n_pages * page_size - BLOB_HEADER_SIZE
            f.seek(skip, 1)
            off += blob.n_pages * page_size
            n += 1
        if skipped_runs:
            pages = sum(count for _start, count in skipped_runs)
            runs = ", ".join(f"{count} at offset {start}" for start, count in skipped_runs[:5])
            more = f" and {len(skipped_runs) - 5} more run(s)" if len(skipped_runs) > 5 else ""
            _warn(
                f"{path}: skipped {pages} page(s) of the blob region that do not "
                f"start a blob ({runs}{more}) and resumed the chain walk after "
                f"each -- leftover or never-written pages between blobs"
            )
        if n > 0 and off > size:
            # The last blob successfully parsed claimed a page count that
            # implies more data than the file actually contains -- off
            # jumped past true EOF. A real, distinct anomaly: the file is
            # cut off in the middle of what should have been that blob's
            # data (or its padding).
            _warn(
                f"{path}: after {n} blob(s), the last one (offset "
                f"{off - blob.n_pages * page_size}, blob_index={blob.blob_index}, "
                f"n_pages={blob.n_pages}) claims data extending "
                f"{off - size} byte(s) past the true end of file ({size} "
                f"byte(s) total) -- file is truncated mid-blob; returning "
                f"the {n} blob header(s) already yielded (note: that last "
                f"blob's own data may itself be incomplete -- see "
                f"read_blob_values()'s truncation handling)"
            )
        elif n > 0 and off != size:
            # Loop condition failed (off + 48 > size) but we're not exactly
            # at the true end either -- real leftover bytes, less than one
            # full header's worth. Every real file checked in this project
            # (docs/provenance/notes.md section 6.6b/6.9) ends with an EXACT match, so any
            # slack here is new/unusual and worth flagging, not silently
            # accepted.
            _warn(
                f"{path}: blob chain walk stopped {size - off} byte(s) short "
                f"of the true end of file (at offset {off} of {size}) after "
                f"{n} blob(s) -- less than one full header remains there, "
                f"which doesn't match any real file checked in this project "
                f"so far (they all end with an exact match); file may be "
                f"truncated"
            )

pygdb.find_blob(path, line_slot, channel_slot, chans_max=None)

Locate the blob for a specific (line, channel) pair.

Parameters:

Name Type Description Default
path str

Path to the .gdb file.

required
line_slot int

0-based physical line-table slot index.

required
channel_slot int

0-based physical channel-table slot index.

required
chans_max int

The file's channel-table capacity. If not given, it's read from the file's own header.

None

Returns:

Type Description
BlobHeader or None

The matching blob header, or None if the chain ends (or breaks -- see iter_blobs's warnings for why) before the target is found, or if chans_max can't be determined at all (bad magic / truncated header) -- never raises for these, consistent with the rest of this module.

Notes

Walks the chain (see iter_blobs) and computes the target blob_index via the formula [CONFIRMED] in docs/provenance/notes.md section 6.6.

This does a linear walk from the start of the blob region every call -- fine for occasional lookups or for building a full line/channel -> offset index once (walk the whole chain yourself with iter_blobs() and record every blob.offset keyed by blob.line_channel(chans_max) if you need many lookups).

Source code in pygdb/gdb_reader.py
def find_blob(path: str, line_slot: int, channel_slot: int, chans_max: Optional[int] = None) -> Optional[BlobHeader]:
    """
    Locate the blob for a specific (line, channel) pair.

    Parameters
    ----------
    path : str
        Path to the `.gdb` file.
    line_slot : int
        0-based physical line-table slot index.
    channel_slot : int
        0-based physical channel-table slot index.
    chans_max : int, optional
        The file's channel-table capacity. If not given, it's read
        from the file's own header.

    Returns
    -------
    BlobHeader or None
        The matching blob header, or `None` if the chain ends (or
        breaks -- see `iter_blobs`'s warnings for why) before the
        target is found, or if `chans_max` can't be determined at all
        (bad magic / truncated header) -- never raises for these,
        consistent with the rest of this module.

    Notes
    -----
    Walks the chain (see `iter_blobs`) and computes the target
    `blob_index` via the formula **[CONFIRMED]** in
    docs/provenance/notes.md section 6.6.

    This does a linear walk from the start of the blob region every
    call -- fine for occasional lookups or for building a full
    line/channel -> offset index once (walk the whole chain yourself
    with `iter_blobs()` and record every `blob.offset` keyed by
    `blob.line_channel(chans_max)` if you need many lookups).
    """
    if chans_max is None:
        with open(path, "rb") as f:
            header = f.read(128)
        if not check_magic(header):
            _warn(f"{path}: does not start with the expected '!CBD' magic -- "
                  f"cannot determine chans_max, blob not found")
            return None
        try:
            chans_max = struct.unpack_from("<i", header, 24)[0]
        except struct.error:
            _warn(f"{path}: header too short to read chans_max -- blob not found")
            return None
    target = line_slot * chans_max + channel_slot
    for blob in iter_blobs(path):
        if blob.blob_index == target:
            return blob
    return None

pygdb.read_blob_directory(path)

Read the persisted blob directory.

Parameters:

Name Type Description Default
path str

Path to the .gdb file.

required

Returns:

Type Description
BlobDirectory or None

The directory, or None when the file has none to trust: a bad magic or truncated header, header words that are inconsistent with the layout (data_slots != lines_max * chans_max, or the data slots would run past the blob region), or a directory whose every data slot is zero. Absence is not an anomaly, so nothing is warned.

Notes

See BlobDirectory. Only the data slots (0 <= slot < data_slots) are read; the registry blob-symbol slots and the cache slots after them are not used by the reader.

Source code in pygdb/gdb_reader.py
def read_blob_directory(path: str) -> Optional[BlobDirectory]:
    """
    Read the persisted blob directory.

    Parameters
    ----------
    path : str
        Path to the `.gdb` file.

    Returns
    -------
    BlobDirectory or None
        The directory, or `None` when the file has none to trust: a bad
        magic or truncated header, header words that are inconsistent
        with the layout (`data_slots != lines_max * chans_max`, or the
        data slots would run past the blob region), or a directory
        whose every data slot is zero. Absence is not an anomaly, so
        nothing is warned.

    Notes
    -----
    See `BlobDirectory`. Only the data slots (`0 <= slot <
    data_slots`) are read; the registry blob-symbol slots and the
    cache slots after them are not used by the reader.
    """
    with open(path, "rb") as f:
        header = f.read(4096)
        if not check_magic(header):
            return None
        fields = header_fields(header)
        chans_max, lines_max, data_slots, page_size = (
            fields["chans_max"], fields["lines_max"], fields["data_slots"], fields["page_size"],
        )
        if None in (chans_max, lines_max, data_slots, page_size):
            return None
        blob_start = blob_region_start(header)
        if (
            blob_start is None or data_slots <= 0 or data_slots != lines_max * chans_max
            or DIRECTORY_OFFSET + data_slots * _DIRECTORY_SLOT_DTYPE.itemsize > blob_start
        ):
            return None
        f.seek(DIRECTORY_OFFSET)
        raw = f.read(data_slots * _DIRECTORY_SLOT_DTYPE.itemsize)
    if len(raw) != data_slots * _DIRECTORY_SLOT_DTYPE.itemsize:
        return None
    slots = np.frombuffer(raw, dtype=_DIRECTORY_SLOT_DTYPE)
    nonzero = np.nonzero((slots["word"] != 0) | (slots["n_pages"] != 0))[0]
    if len(nonzero) == 0:
        return None
    entries = {
        int(i): (int(slots["word"][i]), int(slots["n_pages"][i])) for i in nonzero
    }
    return BlobDirectory(data_slots=data_slots, entries=entries)

pygdb.read_blob_symbols(path)

Read the names of the file's live administrative objects.

Parameters:

Name Type Description Default
path str

Path to the .gdb file.

required

Returns:

Type Description
dict of {int : str} or None

{symbol_slot: name} for every live blob symbol. The object named by slot k is the administrative blob whose blob_index is data_slots + k (header_fields()["data_slots"]). None when the file has no table to trust: a bad magic or truncated header, or header words that would put the table outside the region before the first blob. Absence is not an anomaly, so nothing is warned.

Notes

[CONFIRMED] layout (docs/spec.md section 2.1, docs/provenance/notes.md section 6.2d): blobs_max 128-byte records starting right after the blob directory, at 280 + 6 * index_slots, each beginning with its NUL-terminated name. A record is live when its category (record +76) is DB_CATEGORY_BLOB_NORMAL (0), which held on exactly the 1,267 corpus records that own an administrative blob. A freed slot has bit 0x10000 set and may still hold its old name, so the category, not the name, decides.

Names seen on real files: four fixed objects ("__dbreg", "Display List", "Line Selection", "Database Extension Objects"), "?|IPJ_<X>:<Y>" projection objects, and "__<n>" REG objects where n is a global symbol handle (blobs_max + lines_max + channel_slot for a channel).

Source code in pygdb/gdb_reader.py
def read_blob_symbols(path: str) -> Optional[Dict[int, str]]:
    """
    Read the names of the file's live administrative objects.

    Parameters
    ----------
    path : str
        Path to the `.gdb` file.

    Returns
    -------
    dict of {int : str} or None
        `{symbol_slot: name}` for every live blob symbol. The object
        named by slot `k` is the administrative blob whose `blob_index`
        is `data_slots + k` (`header_fields()["data_slots"]`). `None`
        when the file has no table to trust: a bad magic or truncated
        header, or header words that would put the table outside the
        region before the first blob. Absence is not an anomaly, so
        nothing is warned.

    Notes
    -----
    **[CONFIRMED]** layout (docs/spec.md section 2.1,
    docs/provenance/notes.md section 6.2d): `blobs_max` 128-byte records
    starting right after the blob directory, at `280 + 6 * index_slots`,
    each beginning with its NUL-terminated name. A record is live when
    its category (record `+76`) is `DB_CATEGORY_BLOB_NORMAL` (0), which
    held on exactly the 1,267 corpus records that own an administrative
    blob. A freed slot has bit `0x10000` set and may still hold its old
    name, so the category, not the name, decides.

    Names seen on real files: four fixed objects (`"__dbreg"`,
    `"Display List"`, `"Line Selection"`, `"Database Extension
    Objects"`), `"?|IPJ_<X>:<Y>"` projection objects, and `"__<n>"` REG
    objects where `n` is a global symbol handle (`blobs_max + lines_max
    + channel_slot` for a channel).
    """
    with open(path, "rb") as f:
        header = f.read(4096)
        if not check_magic(header):
            return None
        fields = header_fields(header)
        blobs_max, index_slots = fields["blobs_max"], fields["index_slots"]
        if blobs_max is None or index_slots is None or blobs_max <= 0 or index_slots <= 0:
            return None
        blob_start = blob_region_start(header)
        start = DIRECTORY_OFFSET + index_slots * _DIRECTORY_SLOT_DTYPE.itemsize
        size = blobs_max * SYMBOL_RECORD_SIZE
        if blob_start is None or start + size > blob_start:
            return None
        f.seek(start)
        raw = f.read(size)
    if len(raw) != size:
        return None
    symbols = {}
    for slot in range(blobs_max):
        rec = raw[slot * SYMBOL_RECORD_SIZE:(slot + 1) * SYMBOL_RECORD_SIZE]
        if struct.unpack_from("<i", rec, _BLOB_SYMBOL_CATEGORY_OFFSET)[0] != _DB_CATEGORY_BLOB_NORMAL:
            continue
        name, is_clean = _read_name(rec, 0)
        if name and is_clean:
            symbols[slot] = name
    return symbols

pygdb.read_blob_values(path, blob, channel, comp_level=0, page_size=None, file=None)

Decode a found blob's real row data.

Uses the owning channel's already-known type (from the symbol table, docs/provenance/notes.md section 6.2).

Parameters:

Name Type Description Default
path str

Path to the .gdb file.

required
blob BlobHeader

The blob to decode, as returned by find_blob/iter_blobs.

required
channel ChannelRecord

The owning channel.

required
comp_level int

The file's declared compression level: 0 (DB_COMP_NONE), 1 (DB_COMP_SPEED), or 2 (DB_COMP_SIZE).

0
page_size int

The file's page size, needed to locate a compressed blob's full span for comp_level != 0. If not given, it's read from the file's own header.

None
file BinaryIO

An already-open binary file handle for path, reused instead of opening path fresh -- pass this if you're calling this function many times for the same file (e.g. GDB does, internally). Reopening path on every call is real, measured overhead (~1.7-1.9x slower per call, benchmarked against this project's real sample corpus -- see the Rust-plan's M4 notes); file=None (the default) keeps this function's plain "just give me a path" behavior for every other caller.

None

Returns:

Type Description
ndarray

See _decode_numeric_or_string's docstring for the exact shape/dtype rules -- 1-D ndarray for a scalar channel, 2-D (n_rows, array_width) for a VA/array channel, dtype matching the channel's GS_* type or a fixed-width Unicode dtype for a string channel.

Warns:

Type Description
GDBParseWarning

Per an explicit engineering request, this fails gracefully rather than raising: a negative row_count (a reserved/administrative blob, docs/provenance/notes.md section 6.4/6.9, not real data), a channel type this reader can't decode, a truncated read (file cut off mid-blob), an unrecognized chunk subtype, or a chunk that fails to decompress (corrupt/truncated compressed data, or lzrw1.LZRW1DecodeError) all return an empty ndarray (np.array([])) with this warning describing what went wrong, instead of raising and losing the caller's place in a larger loop (e.g. a whole-file scan that's decoded hundreds of blobs already). The one exception where full graceful salvage wasn't attempted is a truncated/corrupt compressed stream: unlike the plain-data case, there's no simple way to hand back "the first K decoded values" from a partially-decompressed zlib/LZRW1 stream, so those cases warn and return an empty array rather than a partial decode -- documented here rather than silently implied to be as complete as the plain-data truncation handling.

Notes

[CONFIRMED] against real ground truth for GS_DOUBLE data and for fixed-width strings, for DB_COMP_NONE (docs/provenance/notes.md section 6.6: a real fid blob decoded this way reproduces the exact CSV ground-truth value, and a real line-channel blob decodes to the correct real line name repeated once per row).

Also handles compressed blobs (comp_level 1=DB_COMP_SPEED or 2=DB_COMP_SIZE), including multi-page ones -- [CONFIRMED] against real ground truth for both single- and multi-page DB_COMP_SIZE (a real single-page blob_index=0 decodes to the known constant 5027; a real 36-page array-channel blob decodes to LEI_Depth's exact known real depth profile, 0.0, 3.0, 6.3, 9.9, ..., repeated once per station -- both matching docs/provenance/notes.md section 6.⅚.2b's independently- established ground truth exactly) and for both single- and multi-page DB_COMP_SPEED (a real 2-page Northing_AMGz55 blob decodes to sane real coordinates with real rDUMMY sentinels). See docs/provenance/notes.md section 6.6d: a multi-page blob has no per-page re-framing -- just read the full n_pages*page_size span instead of one page. A DB_COMP_SPEED blob is, however, a chain of chunks of at most 16368 decompressed bytes each, not one chunk (docs/provenance/notes.md section 6.6e): only the first carries the 16-byte magic, later ones are a bare 12-byte sub-header plus payload, and the blob header's +24 field is the total decompressed size across the chain. Decoding only the first chunk -- as this function once did -- silently truncated any channel longer than 2046 float64 values on a line.

A real third on-disk variant, auto-detected here rather than assumed away (docs/provenance/notes.md section 6.6b): even inside a file that genuinely declares (and elsewhere uses) DB_COMP_SPEED, some individual blobs turn out to carry no chunk wrapper at all -- just the plain 48-byte DB_COMP_NONE-style header with real, directly readable data straight after it (confirmed on a real Easting_AMGz55 blob in DB_EM_293.gdb: decoding it as if comp_level==0 reproduces sane, real coordinate values with real rDUMMY=-1.0E32 sentinels in the expected places). When comp_level != 0, this function checks for the 16-byte chunk magic at the 56-byte-header position first and only falls back to the genuinely-compressed path if it's actually there -- otherwise it decodes the blob exactly like a DB_COMP_NONE one.

Source code in pygdb/gdb_reader.py
def read_blob_values(path: str, blob: BlobHeader, channel: ChannelRecord,
                      comp_level: int = 0, page_size: Optional[int] = None,
                      file: Optional[BinaryIO] = None):
    """
    Decode a found blob's real row data.

    Uses the owning channel's already-known type (from the symbol
    table, docs/provenance/notes.md section 6.2).

    Parameters
    ----------
    path : str
        Path to the `.gdb` file.
    blob : BlobHeader
        The blob to decode, as returned by `find_blob`/`iter_blobs`.
    channel : ChannelRecord
        The owning channel.
    comp_level : int, default 0
        The file's declared compression level: 0 (`DB_COMP_NONE`), 1
        (`DB_COMP_SPEED`), or 2 (`DB_COMP_SIZE`).
    page_size : int, optional
        The file's page size, needed to locate a compressed blob's
        full span for `comp_level != 0`. If not given, it's read from
        the file's own header.
    file : BinaryIO, optional
        An already-open binary file handle for `path`, reused instead
        of opening `path` fresh -- pass this if you're calling this
        function many times for the same file (e.g. `GDB` does,
        internally). Reopening `path` on every call is real, measured
        overhead (~1.7-1.9x slower per call, benchmarked against this
        project's real sample corpus -- see the Rust-plan's M4 notes);
        `file=None` (the default) keeps this function's plain "just
        give me a path" behavior for every other caller.

    Returns
    -------
    numpy.ndarray
        See `_decode_numeric_or_string`'s docstring for the exact
        shape/dtype rules -- 1-D `ndarray` for a scalar channel, 2-D
        `(n_rows, array_width)` for a VA/array channel, dtype matching
        the channel's `GS_*` type or a fixed-width Unicode dtype for a
        string channel.

    Warns
    -----
    GDBParseWarning
        Per an explicit engineering request, this fails gracefully
        rather than raising: a negative `row_count` (a
        reserved/administrative blob, docs/provenance/notes.md section
        6.4/6.9, not real data), a channel type this reader can't
        decode, a truncated read (file cut off mid-blob), an
        unrecognized chunk subtype, or a chunk that fails to
        decompress (corrupt/truncated compressed data, or
        `lzrw1.LZRW1DecodeError`) all return an empty `ndarray`
        (`np.array([])`) with this warning describing what went wrong,
        instead of raising and losing the caller's place in a larger
        loop (e.g. a whole-file scan that's decoded hundreds of blobs
        already). The one exception where full graceful salvage
        wasn't attempted is a truncated/corrupt *compressed* stream:
        unlike the plain-data case, there's no simple way to hand back
        "the first K decoded values" from a partially-decompressed
        zlib/LZRW1 stream, so those cases warn and return an empty
        array rather than a partial decode -- documented here rather
        than silently implied to be as complete as the plain-data
        truncation handling.

    Notes
    -----
    **[CONFIRMED]** against real ground truth for GS_DOUBLE data and
    for fixed-width strings, for DB_COMP_NONE (docs/provenance/notes.md
    section 6.6: a real `fid` blob decoded this way reproduces the
    exact CSV ground-truth value, and a real `line`-channel blob
    decodes to the correct real line name repeated once per row).

    Also handles compressed blobs (`comp_level` 1=DB_COMP_SPEED or
    2=DB_COMP_SIZE), **including multi-page ones** -- **[CONFIRMED]**
    against real ground truth for both single- and multi-page
    DB_COMP_SIZE (a real single-page blob_index=0 decodes to the known
    constant 5027; a real 36-page array-channel blob decodes to
    `LEI_Depth`'s exact known real depth profile, `0.0, 3.0, 6.3, 9.9,
    ...`, repeated once per station -- both matching
    docs/provenance/notes.md section 6.5/6.2b's independently-
    established ground truth exactly) and for both single- and
    multi-page DB_COMP_SPEED (a real 2-page `Northing_AMGz55` blob
    decodes to sane real coordinates with real `rDUMMY` sentinels).
    See docs/provenance/notes.md section 6.6d: a multi-page blob has no
    per-page re-framing -- just read the full `n_pages*page_size` span
    instead of one page. **A DB_COMP_SPEED blob is, however, a chain of
    chunks of at most 16368 decompressed bytes each, not one chunk**
    (docs/provenance/notes.md section 6.6e): only the first carries the
    16-byte magic, later ones are a bare 12-byte sub-header plus
    payload, and the blob header's `+24` field is the total
    decompressed size across the chain. Decoding only the first chunk
    -- as this function once did -- silently truncated any channel
    longer than 2046 float64 values on a line.

    **A real third on-disk variant, auto-detected here rather than
    assumed away (docs/provenance/notes.md section 6.6b):** even
    inside a file that genuinely declares (and elsewhere uses)
    DB_COMP_SPEED, some individual blobs turn out to carry no chunk
    wrapper at all -- just the plain 48-byte DB_COMP_NONE-style header
    with real, directly readable data straight after it (confirmed on
    a real `Easting_AMGz55` blob in `DB_EM_293.gdb`: decoding it as if
    `comp_level==0` reproduces sane, real coordinate values with real
    `rDUMMY=-1.0E32` sentinels in the expected places). When
    `comp_level != 0`, this function checks for the 16-byte chunk
    magic at the 56-byte-header position first and only falls back to
    the genuinely-compressed path if it's actually there -- otherwise
    it decodes the blob exactly like a DB_COMP_NONE one.
    """
    if comp_level == 0:
        if blob.row_count < 0:
            _warn(
                f"blob_index={blob.blob_index}: negative row_count "
                f"({blob.row_count}) -- this is one of the reserved/"
                f"administrative blobs flagged [UNKNOWN] in docs/provenance/notes.md "
                f"section 6.4/6.9, not a real data blob; returning no values"
            )
            return np.array([])
        width = _element_width(channel)
        if width is None:
            _warn(
                f"channel {channel.name!r} has dtype_code={channel.dtype_code}, a "
                f"type this reader doesn't know how to decode -- returning no values"
            )
            return np.array([])
        with _file_handle(path, file) as f:
            f.seek(blob.data_offset)
            raw = _read_writable(f, blob.row_count * width)
        return _decode_numeric_or_string(raw, channel, blob.row_count)

    # comp_level != 0: could still be any of three real on-disk variants
    # (docs/provenance/notes.md section 6.6b) -- check which one this specific blob
    # actually is rather than assuming from the file-level comp_level.
    with _file_handle(path, file) as f:
        f.seek(blob.offset + COMPRESSED_BLOB_HEADER_SIZE)
        chunk_magic_probe = f.read(8)
    if len(chunk_magic_probe) < 8:
        _warn(
            f"blob_index={blob.blob_index}: file ends before the compressed-blob "
            f"header/chunk-magic region (offset {blob.offset + COMPRESSED_BLOB_HEADER_SIZE}) "
            f"could be fully read -- truncated mid-blob; returning no values"
        )
        return np.array([])
    if chunk_magic_probe != _lzrw1.CHUNK_MAGIC:
        # Variant 3: no chunk wrapper at all -- a "bare" blob, byte-for-byte
        # identical in layout to a DB_COMP_NONE one, just living inside an
        # otherwise-compressed file. Use the plain 48-byte-header fields,
        # which decoded sanely for real in this exact case.
        if blob.row_count < 0:
            _warn(
                f"blob_index={blob.blob_index}: negative row_count "
                f"({blob.row_count}) -- this is one of the reserved/"
                f"administrative blobs flagged [UNKNOWN] in docs/provenance/notes.md "
                f"section 6.4/6.9, not a real data blob; returning no values"
            )
            return np.array([])
        width = _element_width(channel)
        if width is None:
            _warn(
                f"channel {channel.name!r} has dtype_code={channel.dtype_code}, a "
                f"type this reader doesn't know how to decode -- returning no values"
            )
            return np.array([])
        with _file_handle(path, file) as f:
            f.seek(blob.data_offset)
            raw = _read_writable(f, blob.row_count * width)
        return _decode_numeric_or_string(raw, channel, blob.row_count)

    # Compressed (DB_COMP_SPEED / DB_COMP_SIZE): 56-byte blob header,
    # then the shared 16-byte page-primitive chunk magic -- see
    # COMPRESSED_BLOB_HEADER_SIZE and docs/provenance/notes.md section 6.6b/6.6d.
    #
    # Multi-page blobs (blob.n_pages > 1) are [CONFIRMED] (section 6.6d)
    # to be a SINGLE continuous compressed stream spanning the whole
    # n_pages*page_size span -- NOT one independently-framed chunk per
    # page. There is no per-page re-framing to handle: reading the full
    # span is sufficient for zlib (zlib.decompressobj() naturally stops
    # at the real end of stream and reports the rest as padding, and its
    # single stream always matches the blob header's `+24` total). LZRW1
    # is different: the span holds a chain of chunks, not one -- see the
    # DB_COMP_SPEED branch below.
    if page_size is None:
        with _file_handle(path, file) as f:
            header = f.read(128)
        try:
            page_size = struct.unpack_from("<i", header, 100)[0]
        except struct.error:
            _warn(
                f"blob_index={blob.blob_index}: header too short to read "
                f"page_size -- cannot decode, returning no values"
            )
            return np.array([])
    with _file_handle(path, file) as f:
        f.seek(blob.offset + COMPRESSED_BLOB_HEADER_SIZE)
        expected_span = blob.n_pages * page_size - COMPRESSED_BLOB_HEADER_SIZE
        raw_span = f.read(expected_span)
    if len(raw_span) < expected_span:
        _warn(
            f"blob_index={blob.blob_index}: expected {expected_span} byte(s) of "
            f"compressed payload but the file only had {len(raw_span)} available "
            f"-- truncated mid-blob; attempting to decode what's there, but this "
            f"may fail or be incomplete"
        )
    if len(raw_span) < 16:
        _warn(
            f"blob_index={blob.blob_index}: not enough bytes to read even the "
            f"chunk sub-header ({len(raw_span)} available, need 16) -- cannot "
            f"decode, returning no values"
        )
        return np.array([])
    subtype = struct.unpack_from("<i", raw_span, 8)[0]
    if subtype == _lzrw1.DB_COMP_SIZE:
        decompressed = None
        # `pygdb._native.zlib_decompress` (flate2 on the `zlib-rs`
        # backend -- see rust/src/lib.rs's docstring) decompresses into
        # a writable `bytearray` with only one internal copy (Rust's
        # own growable decode buffer -> the final Python object) --
        # better than the stdlib fallback below, which pays that same
        # copy PLUS a second one (`bytes` -> `bytearray`), since stdlib
        # `zlib` can only ever hand back immutable `bytes`. Falls
        # through to stdlib on any failure (corrupt/truncated data)
        # rather than trying to be clever about partial output --
        # stdlib's own graceful-degradation warning below covers it.
        if _native_ext is not None:
            try:
                decompressed = _native_ext.zlib_decompress(raw_span[16:])
            except ValueError:
                decompressed = None
        if decompressed is None:
            try:
                d = zlib.decompressobj()
                # `zlib`'s stdlib API has no way to decompress into a
                # caller-supplied buffer -- `decompress()` always hands
                # back a fresh, immutable `bytes` object. Wrapping it in
                # a `bytearray` here is a deliberate, explicit copy so
                # that every numeric channel ends up writable regardless
                # of which compression mode or backend produced it.
                decompressed = bytearray(d.decompress(raw_span[16:]))
            except zlib.error as e:
                _warn(
                    f"blob_index={blob.blob_index}: zlib decompression failed ({e}) "
                    f"-- likely truncated or corrupt compressed data; returning no values"
                )
                return np.array([])
    elif subtype == _lzrw1.DB_COMP_SPEED:
        # A blob is a chain of chunks of at most 16368 decompressed bytes
        # each (docs/spec.md section 7.3); the blob header's `+24` field
        # is the total across all of them, and is what tells the decoder
        # where the chain ends (the bytes after the last chunk are page
        # padding, not a reliable terminator).
        with _file_handle(path, file) as f:
            f.seek(blob.offset + 24)
            total_field = f.read(4)
        total_decompressed = (
            struct.unpack("<i", total_field)[0] if len(total_field) == 4 else 0
        )
        try:
            decompressed = _lzrw1.decode_speed_blob(raw_span, total_decompressed)
        except _lzrw1.LZRW1DecodeError as e:
            _warn(
                f"blob_index={blob.blob_index}: LZRW1 chunk decode failed ({e}) "
                f"-- likely truncated or corrupt compressed data, or an "
                f"unrecognized chunk variant; returning no values"
            )
            return np.array([])
    else:
        _warn(
            f"blob_index={blob.blob_index}: unrecognized chunk subtype={subtype} "
            f"(expected {_lzrw1.DB_COMP_SIZE}=Size or {_lzrw1.DB_COMP_SPEED}=Speed) "
            f"-- returning no values"
        )
        return np.array([])
    return _decode_numeric_or_string(decompressed, channel)

.grd reader

pygdb.read_grd(path)

Read a .grd file.

Parameters:

Name Type Description Default
path str

Path to the .grd file.

required

Returns:

Name Type Description
header GrdHeader

The parsed 512-byte header.

values array

The grid's raw (unscaled) element values in on-disk order (row-major per the file's own ordering/KX flag). numpy is a dependency of the pygdb package as a whole (see gdb_reader.py's VA/array-channel decoding), but this module's own .grd reading doesn't need it -- a grid's shape is already fully known from shape_e/shape_v, so reshaping is left to the caller rather than done here.

Raises:

Type Description
ValueError

If path is too short to even hold the 512-byte header -- there's nothing to return at all in that case.

Warns:

Type Description
GRDParseWarning

If a compressed block is truncated or fails to decompress (see _decompress_body), or if the decoded element count doesn't match shape_e * shape_v -- which for a real, complete file never happens, so it's always a sign of trouble. values is returned exactly as long as what was actually decoded rather than raising; a caller can check len(values) against header.shape_e * header.shape_v itself if it needs to know.

Source code in pygdb/grd_reader.py
def read_grd(path: str):
    """
    Read a `.grd` file.

    Parameters
    ----------
    path : str
        Path to the `.grd` file.

    Returns
    -------
    header : GrdHeader
        The parsed 512-byte header.
    values : array.array
        The grid's raw (unscaled) element values in on-disk order
        (row-major per the file's own `ordering`/KX flag). numpy is a
        dependency of the `pygdb` package as a whole (see
        `gdb_reader.py`'s VA/array-channel decoding), but this
        module's own `.grd` reading doesn't need it -- a grid's shape
        is already fully known from `shape_e`/`shape_v`, so reshaping
        is left to the caller rather than done here.

    Raises
    ------
    ValueError
        If `path` is too short to even hold the 512-byte header --
        there's nothing to return at all in that case.

    Warns
    -----
    GRDParseWarning
        If a compressed block is truncated or fails to decompress
        (see `_decompress_body`), or if the decoded element count
        doesn't match `shape_e * shape_v` -- which for a real,
        complete file never happens, so it's always a sign of trouble.
        `values` is returned exactly as long as what was actually
        decoded rather than raising; a caller can check `len(values)`
        against `header.shape_e * header.shape_v` itself if it needs
        to know.
    """
    with open(path, "rb") as f:
        raw = f.read()

    if len(raw) < HEADER_SIZE:
        raise ValueError(
            f"{path}: only {len(raw)} byte(s) available, need at least "
            f"{HEADER_SIZE} for the header -- nothing to decode"
        )

    header = parse_header(raw[:HEADER_SIZE])
    body = raw[HEADER_SIZE:]

    if header.is_compressed:
        expected_total_size = header.shape_e * header.shape_v * header.element_size
        body = _decompress_body(body, expected_total_size)

    typecode = _array_typecode(header.element_size, header.sign_flag)
    values = array.array(typecode)
    # array.frombytes requires a whole number of elements; a truncated
    # tail (partial last element) is silently dropped rather than raising,
    # consistent with "return everything that was actually decodable."
    usable = len(body) - (len(body) % values.itemsize)
    values.frombytes(body[:usable])

    expected = header.shape_e * header.shape_v
    if len(values) != expected:
        _warn(
            f"{path}: decoded {len(values)} element(s), expected "
            f"shape_e*shape_v={expected} -- file is likely truncated or "
            f"a compressed block failed to decompress (see any earlier "
            f"GRDParseWarning for which); returning the {len(values)} "
            f"element(s) actually decoded"
        )

    return header, values

pygdb.GrdHeader(n_bytes_per_element, sign_flag, shape_e, shape_v, ordering, spacing_e, spacing_v, x_origin, y_origin, rotation, base_value, data_factor, raw) dataclass

element_size property

Bytes per element, with the +1024 compressed marker stripped.

Registry

pygdb.find_coordinate_systems(path, max_real_line_slot=None)

Scan path for coordinate-system (map projection) names.

Parameters:

Name Type Description Default
path str

Path to the .gdb file to scan.

required
max_real_line_slot int

The highest physical line-table slot index that corresponds to a real survey line -- blobs whose line_slot (decoded via BlobHeader.line_channel) beyond this are treated as "administrative" and probed for IPJ content. If not given, it is derived by calling read_lines(path) (an extra table scan) and using the highest slot index found there; pass it explicitly if you already have that file's read_lines() result to avoid repeating the scan.

None

Returns:

Type Description
list of str

A de-duplicated, order-of-discovery list of name strings (e.g. "WGS 84 / UTM zone 54S"), or an empty list if none were found -- which is expected and normal for a real file with no REG/IPJ content at all (docs/spec.md section 9), not necessarily a sign of a problem.

Warns:

Type Description
GDBParseWarning

If path doesn't start with the expected magic, or its header is too short to read chans_max -- fails gracefully like the rest of this package, returning [] rather than raising.

Source code in pygdb/registry.py
def find_coordinate_systems(path: str, max_real_line_slot: Optional[int] = None) -> List[str]:
    """
    Scan `path` for coordinate-system (map projection) names.

    Parameters
    ----------
    path : str
        Path to the `.gdb` file to scan.
    max_real_line_slot : int, optional
        The highest physical line-table slot index that corresponds to
        a real survey line -- blobs whose `line_slot` (decoded via
        `BlobHeader.line_channel`) beyond this are treated as
        "administrative" and probed for IPJ content. If not given, it
        is derived by calling `read_lines(path)` (an extra table scan)
        and using the highest slot index found there; pass it
        explicitly if you already have that file's `read_lines()`
        result to avoid repeating the scan.

    Returns
    -------
    list of str
        A de-duplicated, order-of-discovery list of name strings (e.g.
        `"WGS 84 / UTM zone 54S"`), or an empty list if none were
        found -- which is expected and normal for a real file with no
        REG/IPJ content at all (docs/spec.md section 9), not
        necessarily a sign of a problem.

    Warns
    -----
    GDBParseWarning
        If `path` doesn't start with the expected magic, or its header
        is too short to read `chans_max` -- fails gracefully like the
        rest of this package, returning `[]` rather than raising.
    """
    with open(path, "rb") as f:
        header = f.read(128)
    if not check_magic(header):
        _warn(f"{path}: does not start with the expected '!CBD' magic -- "
              f"no coordinate systems")
        return []
    fields = header_fields(header)
    chans_max = fields["chans_max"]
    if chans_max is None:
        _warn(f"{path}: header too short to read chans_max -- "
              f"no coordinate systems")
        return []

    if max_real_line_slot is None:
        lines = read_lines(path)
        max_real_line_slot = max((line.index for line in lines), default=-1)

    names: List[str] = []
    seen = set()
    with open(path, "rb") as f:
        for blob in iter_blobs(path):
            line_slot, _channel_slot = blob.line_channel(chans_max)
            if line_slot <= max_real_line_slot:
                continue  # a real survey line's data, not administrative metadata
            f.seek(blob.offset)
            chunk = f.read(_PROBE_SIZE)
            m = _IPJ_NAME_RE.search(chunk)
            if not m:
                continue
            name = m.group(1).decode("ascii", errors="replace")
            if name not in seen:
                seen.add(name)
                names.append(name)
    return names

pygdb.find_channel_roles(path, max_real_line_slot=None, channel_names=None)

Scan path for which real channel plays the X/Y/Z coordinate role.

Reads the file's own internal registry (docs/provenance/notes.md section 6.8b) -- a directly-decodable alternative to guessing from channel-naming conventions ("Easting"/"Northing" and similar aren't consistent enough across real files to guess safely; see gdb.GDB.to_xarray's docstring for why this reader avoids that kind of guess elsewhere too).

Parameters:

Name Type Description Default
path str

Path to the .gdb file to scan.

required
max_real_line_slot int

See find_coordinate_systems -- same meaning and same "pass it if you already have it" reasoning.

None
channel_names iterable of str

The file's own real channel names, used to validate each candidate registry value (see Notes). If not given, this calls read_channels(path) itself (an extra table scan) -- pass [c.name for c in db.channels] if the caller already has it.

None

Returns:

Type Description
dict of {str : str or None}

{"X": ..., "Y": ..., "Z": ...}, always all three keys; a role with no confirmed real-channel assignment (the registry key is absent, its value doesn't match any real channel in channel_names, or -- a real, confirmed case -- its value is a single blank space, Oasis montaj's own "no channel assigned to this role" placeholder) maps to None rather than being omitted, so a caller doesn't need to distinguish "not found" from "found but unusable."

Warns:

Type Description
GDBParseWarning

If path doesn't start with the expected magic, or its header is too short to read chans_max/page_size -- fails gracefully like find_coordinate_systems, returning all-None rather than raising. Also raised if more than one different candidate value both validate as real channels for the same role (see Notes) -- genuine ambiguity, not yet observed on any real file.

Notes

A real complication, found by testing (section 6.8b): this format's append-only blob storage can leave multiple, differing stale copies of the same registry key in one file when a role gets re-registered (confirmed on 2 of 22 real files) -- and neither "prefer the first occurrence" nor "prefer the last" resolves both real cases correctly (one needs each). The robust rule used here instead: collect every candidate value found for a role, and keep whichever one(s) actually name a real, current channel (checked against channel_names) -- a direct cross-check against data this reader already parses, not a positional guess.

Source code in pygdb/registry.py
def find_channel_roles(
    path: str,
    max_real_line_slot: Optional[int] = None,
    channel_names: Optional[Iterable[str]] = None,
) -> Dict[str, Optional[str]]:
    """
    Scan `path` for which real channel plays the X/Y/Z coordinate role.

    Reads the file's own internal registry (docs/provenance/notes.md
    section 6.8b) -- a directly-decodable alternative to guessing from
    channel-naming conventions (`"Easting"`/`"Northing"` and similar
    aren't consistent enough across real files to guess safely; see
    `gdb.GDB.to_xarray`'s docstring for why this reader avoids that
    kind of guess elsewhere too).

    Parameters
    ----------
    path : str
        Path to the `.gdb` file to scan.
    max_real_line_slot : int, optional
        See `find_coordinate_systems` -- same meaning and same
        "pass it if you already have it" reasoning.
    channel_names : iterable of str, optional
        The file's own real channel names, used to validate each
        candidate registry value (see Notes). If not given, this calls
        `read_channels(path)` itself (an extra table scan) -- pass
        `[c.name for c in db.channels]` if the caller already has it.

    Returns
    -------
    dict of {str : str or None}
        `{"X": ..., "Y": ..., "Z": ...}`, always all three keys; a role
        with no confirmed real-channel assignment (the registry key is
        absent, its value doesn't match any real channel in
        `channel_names`, or -- a real, confirmed case -- its value is a
        single blank space, Oasis montaj's own "no channel assigned to
        this role" placeholder) maps to `None` rather than being
        omitted, so a caller doesn't need to distinguish "not found"
        from "found but unusable."

    Warns
    -----
    GDBParseWarning
        If `path` doesn't start with the expected magic, or its header
        is too short to read `chans_max`/`page_size` -- fails
        gracefully like `find_coordinate_systems`, returning all-`None`
        rather than raising. Also raised if more than one *different*
        candidate value both validate as real channels for the same
        role (see Notes) -- genuine ambiguity, not yet observed on any
        real file.

    Notes
    -----
    A real complication, found by testing (section 6.8b): this
    format's append-only blob storage can leave *multiple, differing*
    stale copies of the same registry key in one file when a role gets
    re-registered (confirmed on 2 of 22 real files) -- and neither
    "prefer the first occurrence" nor "prefer the last" resolves both
    real cases correctly (one needs each). The robust rule used here
    instead: collect every candidate value found for a role, and keep
    whichever one(s) actually name a real, current channel (checked
    against `channel_names`) -- a direct cross-check against data this
    reader already parses, not a positional guess.
    """
    with open(path, "rb") as f:
        header = f.read(128)
    if not check_magic(header):
        _warn(f"{path}: does not start with the expected '!CBD' magic -- "
              f"no channel roles")
        return {role: None for role in _CHANNEL_ROLE_KEYS}
    fields = header_fields(header)
    chans_max = fields["chans_max"]
    page_size = fields["page_size"]
    if chans_max is None or not page_size:
        _warn(f"{path}: header too short to read chans_max/page_size -- "
              f"no channel roles")
        return {role: None for role in _CHANNEL_ROLE_KEYS}

    if max_real_line_slot is None:
        lines = read_lines(path)
        max_real_line_slot = max((line.index for line in lines), default=-1)

    if channel_names is None:
        channel_names = {c.name for c in read_channels(path)}
    else:
        channel_names = set(channel_names)

    candidates: Dict[str, List[str]] = {role: [] for role in _CHANNEL_ROLE_KEYS}
    with open(path, "rb") as f:
        for blob in iter_blobs(path):
            line_slot, _channel_slot = blob.line_channel(chans_max)
            if line_slot <= max_real_line_slot:
                continue  # a real survey line's data, not administrative metadata
            f.seek(blob.offset)
            # Read the blob's own full declared extent, not a fixed-size
            # probe: unlike `find_coordinate_systems`'s IPJ name marker
            # (confirmed to sit near the start of its blob), a real file
            # was found where `DB_CHAN_X` sits 20096 bytes into a
            # 32768-byte blob -- comfortably past a `_PROBE_SIZE=2000`
            # window, which would silently miss it. Capped well above
            # any real blob size seen in this project's corpus (largest
            # ~16KB decompressed per the Rust-plan notes) purely as a
            # guard against a corrupt/absurd `n_pages` value, not a
            # tuned-to-real-data limit the way `_PROBE_SIZE` is.
            blob_size = min(blob.n_pages * page_size, 50_000_000)
            chunk = f.read(blob_size)
            for role, key in _CHANNEL_ROLE_KEYS.items():
                idx = chunk.find(key)
                if idx == -1:
                    continue
                start = idx + len(key)
                end = chunk.find(b"\x00", start)
                if end == -1:
                    continue  # truncated read -- the value ran past the probe window
                candidates[role].append(chunk[start:end].decode("ascii", errors="replace"))

    roles: Dict[str, Optional[str]] = {}
    for role, values in candidates.items():
        valid = {v for v in values if v in channel_names}
        if len(valid) > 1:
            _warn(
                f"{path}: found {len(valid)} different real channels "
                f"registered for the {role} coordinate role ({sorted(valid)!r}) "
                f"across stale/duplicate registry entries -- ambiguous, not "
                f"using any of them"
            )
            roles[role] = None
        else:
            roles[role] = next(iter(valid), None)
    return roles

pygdb.find_channel_settings(path, channels=None)

Scan path for real per-channel settings recorded in its REG registry.

Reads the same "REG "-tagged administrative-blob content find_coordinate_systems/find_channel_roles already scan, but decodes its flat key/value framing directly (docs/provenance/notes.md section 6.8c) instead of searching for one specific marker -- so this surfaces whatever real settings a channel's REG entries happen to carry: real per-channel display units (UNITS), real user-entered processing labels (LABEL), real processing formulas (FORMULA), among others (docs/spec.md section 9 has the cross-validated evidence for what these keys mean in practice).

Parameters:

Name Type Description Default
path str

Path to the .gdb file to scan.

required
channels iterable of ChannelRecord

The file's own real channel table, used to resolve the channel slot an object's symbol name points at to a real channel name (see Notes). If not given, this calls read_channels(path) itself (an extra table scan) -- pass db.channels if the caller already has it.

None

Returns:

Type Description
dict of {str : dict of {str : str}}

{channel_name: {key: value, ...}, ...} -- only for a channel with at least one populated key found in its own REG entries; a channel with none (no REG entry at all, or entries that are all bare placeholders, or objects holding only a nested MAKER record -- see Notes) is simply absent, not mapped to {}. Also {} for a file with no readable blob-symbol table (gdb_reader.read_blob_symbols returns None), since nothing then says which channel an object belongs to.

Warns:

Type Description
GDBParseWarning

If path doesn't start with the expected magic, or its header is too short to read chans_max/page_size -- fails gracefully like the sibling functions, returning {} rather than raising. Also raised, once per (channel, key) pair, if two or more of a channel's REG entries give different non-empty values for the same key -- the last one in blob-chain order is kept, but this is never decided silently.

Notes

Only an object's KEY\0value\0 entries are decoded (see _decode_reg_flat_keyvalues). Its nested objects -- in practice a MAKER record naming the GX that made the channel (docs/spec.md section 9) -- are not. Which keys are meaningful for a given channel is not itself decoded from anything -- only the keys a real file happens to have written are returned.

Which channel an object belongs to comes from its blob symbol (docs/spec.md section 2.1, docs/provenance/notes.md section 6.2d): the administrative blob at blob_index = data_slots + k is named by blob-symbol slot k, and a per-channel REG object is named "__<n>", where n is the channel's global symbol handle, blobs_max + lines_max + channel_slot. [CONFIRMED]: where an object's LABEL equals a real channel's name, this mapping names that channel on 218 corpus objects. The blob index's own remainder (blob_index % chans_max), which an earlier version of this function used, named it on none. An object whose handle is not a channel's -- a line handle (real, 744 corpus objects, meaning [UNKNOWN]) or any other name -- is skipped.

This format's append-only storage can leave stale, differing copies of the same object (the same phenomenon find_channel_roles handles for DB_CHAN_X/Y/Z and issue #2's data blobs) -- every occurrence found is collapsed per key, last one in blob-chain order wins, with a warning only when two real occurrences actually disagree.

Source code in pygdb/registry.py
def find_channel_settings(
    path: str,
    channels: Optional[Iterable[ChannelRecord]] = None,
) -> Dict[str, Dict[str, str]]:
    """
    Scan `path` for real per-channel settings recorded in its REG registry.

    Reads the same "REG "-tagged administrative-blob content
    `find_coordinate_systems`/`find_channel_roles` already scan, but
    decodes its flat key/value framing directly (docs/provenance/notes.md
    section 6.8c) instead of searching for one specific marker -- so this
    surfaces whatever real settings a channel's REG entries happen to
    carry: real per-channel display units (`UNITS`), real user-entered
    processing labels (`LABEL`), real processing formulas (`FORMULA`),
    among others (docs/spec.md section 9 has the cross-validated evidence
    for what these keys mean in practice).

    Parameters
    ----------
    path : str
        Path to the `.gdb` file to scan.
    channels : iterable of ChannelRecord, optional
        The file's own real channel table, used to resolve the channel
        slot an object's symbol name points at to a real channel name
        (see Notes). If not given, this calls `read_channels(path)`
        itself (an extra table scan) -- pass `db.channels` if the caller
        already has it.

    Returns
    -------
    dict of {str : dict of {str : str}}
        `{channel_name: {key: value, ...}, ...}` -- only for a channel
        with at least one populated key found in its own REG entries; a
        channel with none (no REG entry at all, or entries that are all
        bare placeholders, or objects holding only a nested `MAKER`
        record -- see Notes) is simply absent, not mapped to `{}`. Also `{}` for a
        file with no readable blob-symbol table
        (`gdb_reader.read_blob_symbols` returns `None`), since nothing
        then says which channel an object belongs to.

    Warns
    -----
    GDBParseWarning
        If `path` doesn't start with the expected magic, or its header
        is too short to read `chans_max`/`page_size` -- fails gracefully
        like the sibling functions, returning `{}` rather than raising.
        Also raised, once per `(channel, key)` pair, if two or more of a
        channel's REG entries give *different* non-empty values for the
        same key -- the last one in blob-chain order is kept, but this is
        never decided silently.

    Notes
    -----
    **Only an object's `KEY\\0value\\0` entries are decoded** (see
    `_decode_reg_flat_keyvalues`). Its nested objects -- in practice a
    `MAKER` record naming the GX that made the channel (docs/spec.md
    section 9) -- are not. Which keys are meaningful for a given channel
    is not itself decoded from anything -- only the keys a real file
    happens to have written are returned.

    **Which channel an object belongs to** comes from its blob symbol
    (docs/spec.md section 2.1, docs/provenance/notes.md section 6.2d):
    the administrative blob at `blob_index = data_slots + k` is named by
    blob-symbol slot `k`, and a per-channel REG object is named
    `"__<n>"`, where `n` is the channel's global symbol handle,
    `blobs_max + lines_max + channel_slot`. **[CONFIRMED]**: where an
    object's `LABEL` equals a real channel's name, this mapping names
    that channel on 218 corpus objects. The blob index's own remainder
    (`blob_index % chans_max`), which an earlier version of this
    function used, named it on none. An object whose handle is not a
    channel's -- a line handle (real, 744 corpus objects, meaning
    **[UNKNOWN]**) or any other name -- is skipped.

    This format's append-only storage can leave stale, differing copies
    of the same object (the same phenomenon `find_channel_roles` handles
    for `DB_CHAN_X/Y/Z` and issue #2's data blobs) -- every occurrence
    found is collapsed per key, last one in blob-chain order wins, with
    a warning only when two real occurrences actually disagree.
    """
    with open(path, "rb") as f:
        header = f.read(128)
    if not check_magic(header):
        _warn(f"{path}: does not start with the expected '!CBD' magic -- "
              f"no channel settings")
        return {}
    fields = header_fields(header)
    page_size = fields["page_size"]
    if None in (fields["chans_max"], fields["blobs_max"], fields["lines_max"],
                fields["data_slots"]) or not page_size:
        _warn(f"{path}: header too short to read the table sizes/page_size -- "
              f"no channel settings")
        return {}
    data_slots = fields["data_slots"]
    channel_handle_base = fields["blobs_max"] + fields["lines_max"]

    symbols = read_blob_symbols(path)
    if symbols is None:
        return {}

    if channels is None:
        channels = read_channels(path)
    channel_names = {c.index: c.name for c in channels}

    # {channel_slot: {key: [value, ...]}}, in blob-chain order.
    candidates: Dict[int, Dict[str, List[str]]] = {}
    with open(path, "rb") as f:
        for blob in iter_blobs(path):
            if blob.blob_index < data_slots:
                continue  # a real survey line's data, not administrative metadata
            m = _CHANNEL_OBJECT_NAME_RE.fullmatch(symbols.get(blob.blob_index - data_slots, ""))
            if not m:
                continue  # not a per-symbol REG object
            channel_slot = int(m.group(1)) - channel_handle_base
            if channel_slot not in channel_names:
                continue  # not a real, current channel's handle -- nothing to attach this to
            f.seek(blob.offset)
            # Same "read the blob's own full declared extent, capped" approach
            # as find_channel_roles, for the same reason: a real key has been
            # found tens of kilobytes into a real blob, well past any small
            # fixed-size probe.
            blob_size = min(blob.n_pages * page_size, 50_000_000)
            chunk = f.read(blob_size)
            kv = _decode_reg_flat_keyvalues(chunk)
            if not kv:
                continue
            by_key = candidates.setdefault(channel_slot, {})
            for key, value in kv.items():
                by_key.setdefault(key, []).append(value)

    settings: Dict[str, Dict[str, str]] = {}
    for channel_slot, by_key in candidates.items():
        name = channel_names[channel_slot]
        resolved: Dict[str, str] = {}
        for key, values in by_key.items():
            distinct = set(values)
            if len(distinct) > 1:
                _warn(
                    f"{path}: channel {name!r} has {len(distinct)} different "
                    f"values for registry key {key!r} across stale/duplicate "
                    f"entries ({sorted(distinct)!r}) -- using the last one in "
                    f"blob-chain order"
                )
            resolved[key] = values[-1]
        if resolved:
            settings[name] = resolved
    return settings

pygdb.find_projection_parameters(path, max_real_line_slot=None)

Scan path for real geodetic parameters recorded in its IPJ registry.

Reads the same IPJ-tagged administrative-blob content find_coordinate_systems already scans for a name, but decodes the object's fixed-offset binary record directly (docs/provenance/ notes.md section 6.7b) instead of stopping at the name -- real datum, ellipsoid, and projection parameters, cross-validated against independent ground truth on real files (docs/spec.md section 8).

Parameters:

Name Type Description Default
path str

Path to the .gdb file to scan.

required
max_real_line_slot int

See find_coordinate_systems -- same meaning and same "pass it if you already have it" reasoning.

None

Returns:

Type Description
dict of {str : ProjectionParameters}

Keyed by the same working coordinate-system name find_coordinate_systems returns; a real file with no IPJ content at all (docs/spec.md section 9) gives {}, which is expected and normal, not a sign of a problem.

Warns:

Type Description
GDBParseWarning

If path doesn't start with the expected magic, or its header is too short to read chans_max -- fails gracefully like the sibling functions, returning {} rather than raising. Also raised, once per name, if two of a coordinate system's IPJ entries decode to genuinely different parameter sets -- the last one in blob-chain order is kept, but this is never decided silently (the same convention find_channel_settings uses). Also raised, once per name, if the registry's _PJ_PROJECTION text for a coordinate system doesn't match its binary parameters; the text is then not used (see ProjectionParameters).

Notes

A candidate blob is only decoded if it has the confirmed IPJ object shape: the type-code field reading b"IPJ\x00", at least 652 bytes (through the last parameter slot), and the " JPI" marker's own tag also present at its fixed +96 position -- a structural gate before trusting the fixed-offset fields, matching the validation find_channel_settings applies to REG objects. Anything else is silently skipped, not warned about.

Source code in pygdb/registry.py
def find_projection_parameters(
    path: str, max_real_line_slot: Optional[int] = None,
) -> Dict[str, ProjectionParameters]:
    """
    Scan `path` for real geodetic parameters recorded in its IPJ registry.

    Reads the same `IPJ`-tagged administrative-blob content
    `find_coordinate_systems` already scans for a name, but decodes the
    object's fixed-offset binary record directly (docs/provenance/
    notes.md section 6.7b) instead of stopping at the name -- real datum,
    ellipsoid, and projection parameters, cross-validated against
    independent ground truth on real files (docs/spec.md section 8).

    Parameters
    ----------
    path : str
        Path to the `.gdb` file to scan.
    max_real_line_slot : int, optional
        See `find_coordinate_systems` -- same meaning and same
        "pass it if you already have it" reasoning.

    Returns
    -------
    dict of {str : ProjectionParameters}
        Keyed by the same working coordinate-system name
        `find_coordinate_systems` returns; a real file with no `IPJ`
        content at all (docs/spec.md section 9) gives `{}`, which is
        expected and normal, not a sign of a problem.

    Warns
    -----
    GDBParseWarning
        If `path` doesn't start with the expected magic, or its header
        is too short to read `chans_max` -- fails gracefully like the
        sibling functions, returning `{}` rather than raising. Also
        raised, once per name, if two of a coordinate system's `IPJ`
        entries decode to genuinely *different* parameter sets -- the
        last one in blob-chain order is kept, but this is never decided
        silently (the same convention `find_channel_settings` uses). Also
        raised, once per name, if the registry's `_PJ_PROJECTION` text
        for a coordinate system doesn't match its binary parameters; the
        text is then not used (see `ProjectionParameters`).

    Notes
    -----
    A candidate blob is only decoded if it has the confirmed `IPJ`
    object shape: the type-code field reading `b"IPJ\\x00"`, at least
    652 bytes (through the last parameter slot), and the `" JPI"`
    marker's own tag also present at its fixed `+96` position -- a
    structural gate before trusting the fixed-offset fields, matching
    the validation `find_channel_settings` applies to `REG` objects.
    Anything else is silently skipped, not warned about.
    """
    with open(path, "rb") as f:
        header = f.read(128)
    if not check_magic(header):
        _warn(f"{path}: does not start with the expected '!CBD' magic -- "
              f"no projection parameters")
        return {}
    fields = header_fields(header)
    chans_max = fields["chans_max"]
    page_size = fields["page_size"]
    if chans_max is None or not page_size:
        _warn(f"{path}: header too short to read chans_max/page_size -- "
              f"no projection parameters")
        return {}

    if max_real_line_slot is None:
        lines = read_lines(path)
        max_real_line_slot = max((line.index for line in lines), default=-1)

    candidates: Dict[str, List[ProjectionParameters]] = {}
    # {coordinate-system name: [_PJ_PROJECTION text, ...]}, in blob-chain order.
    texts: Dict[str, List[str]] = {}
    with open(path, "rb") as f:
        for blob in iter_blobs(path):
            line_slot, _channel_slot = blob.line_channel(chans_max)
            if line_slot <= max_real_line_slot:
                continue  # a real survey line's data, not administrative metadata
            f.seek(blob.offset)
            blob_size = min(blob.n_pages * page_size, 50_000_000)
            chunk = f.read(blob_size)
            kv = _decode_reg_flat_keyvalues(chunk)
            if kv and "_PJ_NAME" in kv and "_PJ_PROJECTION" in kv:
                texts.setdefault(kv["_PJ_NAME"].strip('"'), []).append(kv["_PJ_PROJECTION"])
                continue
            if (
                len(chunk) < _IPJ_MIN_LENGTH
                or chunk[96:100] != _IPJ_GATE_TAG
            ):
                continue
            m = _IPJ_NAME_RE.search(chunk)
            if not m:
                continue
            name = m.group(1).decode("ascii", errors="replace")
            datum_name = _read_ascii_cstr(chunk, _IPJ_DATUM_NAME_OFFSET)
            ellipsoid_name = _read_ascii_cstr(chunk, _IPJ_ELLIPSOID_NAME_OFFSET)
            if datum_name is None or ellipsoid_name is None:
                continue  # truncated read -- a name ran past what was read
            transform_name = _read_ascii_cstr(chunk, _IPJ_DATUM_TRANSFORM_NAME_OFFSET)
            if transform_name is not None and " to " not in transform_name:
                transform_name = None  # the datum's own name repeated, not a real transform

            raw = struct.unpack_from(f"<{_IPJ_PARAMETER_COUNT}d", chunk, _IPJ_PARAMETERS_OFFSET)
            slots = tuple(None if v == _IPJ_DUMMY_FLOAT else v for v in raw)
            method = struct.unpack_from("<i", chunk, _IPJ_METHOD_OFFSET)[0]
            layout = _IPJ_SLOTS.get(method, {})
            named = {key: slots[slot] for key, slot in layout.items()}

            def _double(offset: int) -> Optional[float]:
                v = struct.unpack_from("<d", chunk, offset)[0]
                return None if v == _IPJ_DUMMY_FLOAT else v

            transform = struct.unpack_from("<7d", chunk, _IPJ_DATUM_TRANSFORM_OFFSET)
            if _IPJ_DUMMY_FLOAT in transform:
                transform_gxf = None
            else:
                transform_gxf = (
                    transform[:3]
                    + tuple(v * _RADIANS_TO_ARC_SECONDS for v in transform[3:6])
                    + ((transform[6] - 1.0) * 1e6,)
                )

            params = ProjectionParameters(
                name=name,
                datum_name=datum_name,
                ellipsoid_name=ellipsoid_name,
                datum_transform_name=transform_name,
                semi_major_axis=struct.unpack_from("<d", chunk, _IPJ_SEMI_MAJOR_AXIS_OFFSET)[0],
                eccentricity=struct.unpack_from("<d", chunk, _IPJ_ECCENTRICITY_OFFSET)[0],
                central_meridian=named.get("central_meridian"),
                scale_factor=named.get("scale_factor"),
                false_easting=named.get("false_easting"),
                false_northing=named.get("false_northing"),
                method_code=method,
                latitude_of_origin=named.get("latitude_of_origin"),
                standard_parallel_1=named.get("standard_parallel_1"),
                standard_parallel_2=named.get("standard_parallel_2"),
                parameters=slots,
                prime_meridian=_double(_IPJ_PRIME_MERIDIAN_OFFSET),
                datum_transform_parameters=transform_gxf,
                units_name=_read_ascii_cstr(chunk, _IPJ_UNITS_NAME_OFFSET) or None,
                units_factor=_double(_IPJ_UNITS_FACTOR_OFFSET),
                projection_name=_read_ascii_cstr(chunk, _IPJ_PROJECTION_NAME_OFFSET) or None,
            )
            candidates.setdefault(name, []).append(params)

    disagreeing_text = set()
    for name, params_list in candidates.items():
        for i, params in enumerate(params_list):
            params_list[i], disagreed = _name_method_parameters(params, texts.get(name, []))
        if disagreed:  # only the copy that is kept (the last) matters
            disagreeing_text.add(name)

    for name in sorted(disagreeing_text):
        _warn(
            f"{path}: registry _PJ_PROJECTION text for coordinate system "
            f"{name!r} does not match its IPJ binary parameters -- not "
            f"using the text"
        )

    result: Dict[str, ProjectionParameters] = {}
    for name, params_list in candidates.items():
        distinct = []
        for p in params_list:
            if p not in distinct:
                distinct.append(p)
        if len(distinct) > 1:
            _warn(
                f"{path}: coordinate system {name!r} has {len(distinct)} "
                f"different sets of projection parameters across "
                f"stale/duplicate IPJ entries -- using the last one in "
                f"blob-chain order"
            )
        result[name] = params_list[-1]
    return result

pygdb.ProjectionParameters(name, datum_name, ellipsoid_name, datum_transform_name, semi_major_axis, eccentricity, central_meridian, scale_factor, false_easting, false_northing, method_code=None, latitude_of_origin=None, standard_parallel_1=None, standard_parallel_2=None, parameters=(), method=None, method_parameters=dict(), parameter_source=None, prime_meridian=None, datum_transform_parameters=None, units_name=None, units_factor=None, projection_name=None) dataclass

Real geodetic parameters decoded from one IPJ registry object.

Attributes:

Name Type Description
name str

The working coordinate-system name (the same string find_coordinate_systems returns for this object).

datum_name str

E.g. "NAD83", "GDA2020", "WGS 84".

ellipsoid_name str

E.g. "GRS 1980", "WGS 84".

datum_transform_name str or None

E.g. "NAD83 to WGS 84 (1)". None when this object's datum is already WGS 84 -- nothing to transform, not a decode failure.

semi_major_axis float

The ellipsoid's semi-major axis, in metres.

eccentricity float

The ellipsoid's eccentricity.

central_meridian, scale_factor, false_easting, false_northing float or None

The projection's own parameters. None when the object defines only a datum/ellipsoid, when the projection method does not use that parameter (Lambert has no scale factor), or when the method is not one this reader knows (see Notes).

method_code int or None

The projection-method code at +168: 1 geographic (datum only), 11 Transverse Mercator, 3 Lambert Conic Conformal (2SP), 14 Polar Stereographic. Other codes are real, but their slot layout isn't known; see method_parameters.

latitude_of_origin float or None

Transverse Mercator or Lambert latitude of origin; for Polar Stereographic, the latitude in the same slot, which the source survey's own metadata calls the standard parallel.

standard_parallel_1, standard_parallel_2 float or None

Lambert Conic Conformal (2SP) standard parallels.

parameters tuple of (float or None)

All eight raw parameter slots (+588..+651) in order, rDUMMY mapped to None.

method str or None

The projection method's GXF name, e.g. "Transverse Mercator", "Geographic". None when neither the method code nor the file's own registry text identifies it.

method_parameters dict of {str : float or None}

The method's parameters keyed by their GXF Table 1 names (lowercased, underscores), e.g. latitude_of_natural_origin, scale_factor_at_natural_origin, false_easting, in Table 1 order. {} for a geographic system, or when the parameters can't be named (parameter_source is None).

parameter_source {'text', 'binary'} or None

Where the names in method_parameters came from (see Notes). "text": the file's own _PJ_PROJECTION text, checked against the binary. "binary": this reader's confirmed slot layout for the method code. None: neither is available, and only parameters holds the values.

prime_meridian float or None

Degrees from Greenwich. None if unset.

datum_transform_parameters tuple of float or None

The 7-parameter Bursa-Wolf datum transform to WGS 84, in GXF units: dX, dY, dZ in metres, Rx, Ry, Rz in arc-seconds, scale in ppm. All zero for a datum that is already WGS 84. None if any value is the unset sentinel.

units_name str or None

The coordinate units, e.g. "m", or "dega" (degrees) for a geographic system.

units_factor float or None

The units' factor to metres. None if unset.

projection_name str or None

The projection's name without its datum, e.g. "UTM zone 11N". None for a geographic system.

Notes

[CONFIRMED] structure and slot positions for the method codes above (docs/spec.md section 8, docs/provenance/notes.md section 6.7b), each against the file's own _PJ_PROJECTION text, with the slot names from Geosoft's GXF Revision 3 specification, Table 1. The on-disk value of an unset parameter is the vendor's float64 rDUMMY sentinel (-1.0e32, docs/spec.md section 4), mapped to None here rather than returned as a raw dummy a caller could mistake for a real coordinate -- this reader's own convention.

Methods without a known code. A registry's _PJ_PROJECTION text ("Method name",p1,p2,...) lists a method's parameters in GXF Table 1 order, leaving out the unused ones; the binary keeps those as unset slots. So the text names the binary's set slots, in order, for any method in Table 1. The text is used only when its values equal the binary's set slots exactly (43 of 43 corpus objects with text) -- the values returned are always the binary's. Text whose method isn't in Table 1, or has a different parameter count, gives method but no names.

The existing named fields (central_meridian, scale_factor, false_easting, false_northing, latitude_of_origin, standard_parallel_1, standard_parallel_2) are filled only for the method codes with a confirmed slot layout, as before.

pygdb.find_channel_makers(path, channels=None)

Scan path for how each channel was made.

Parameters:

Name Type Description Default
path str

Path to the .gdb file to scan.

required
channels iterable of ChannelRecord

The file's channel table, used to resolve each record's owning channel handle to a name. If not given, this calls read_channels(path) itself -- pass db.channels if the caller already has it.

None

Returns:

Type Description
dict of {str : ChannelMaker}

{channel name: ChannelMaker} for every channel whose registry object holds a MAKER record. A channel with none -- imported rather than made by a tool, or a file whose writer kept no record -- is absent. {} for a file with no readable blob-symbol table.

Warns:

Type Description
GDBParseWarning

If path doesn't start with the expected magic. Also raised, once per channel, if stale copies of its registry object hold different MAKER records -- the last one in blob-chain order is kept.

Notes

Attribution follows find_channel_settings: the record is owned by the channel named by its object's blob-symbol handle (docs/spec.md section 2.1). Every copy of the object in the chain is read, since a rewritten object can leave its MAKER record only in an older copy of itself; differing copies are warned about, never merged.

Source code in pygdb/registry.py
def find_channel_makers(
    path: str,
    channels: Optional[Iterable[ChannelRecord]] = None,
) -> Dict[str, ChannelMaker]:
    """
    Scan `path` for how each channel was made.

    Parameters
    ----------
    path : str
        Path to the `.gdb` file to scan.
    channels : iterable of ChannelRecord, optional
        The file's channel table, used to resolve each record's owning
        channel handle to a name. If not given, this calls
        `read_channels(path)` itself -- pass `db.channels` if the caller
        already has it.

    Returns
    -------
    dict of {str : ChannelMaker}
        `{channel name: ChannelMaker}` for every channel whose registry
        object holds a `MAKER` record. A channel with none -- imported
        rather than made by a tool, or a file whose writer kept no record
        -- is absent. `{}` for a file with no readable blob-symbol table.

    Warns
    -----
    GDBParseWarning
        If `path` doesn't start with the expected magic. Also raised, once
        per channel, if stale copies of its registry object hold different
        `MAKER` records -- the last one in blob-chain order is kept.

    Notes
    -----
    Attribution follows `find_channel_settings`: the record is owned by
    the channel named by its object's blob-symbol handle (docs/spec.md
    section 2.1). Every copy of the object in the chain is read, since a
    rewritten object can leave its `MAKER` record only in an older copy
    of itself; differing copies are warned about, never merged.
    """
    resolved = _channel_handles(path, channels)
    if resolved is None:
        _warn(f"{path}: not a readable .gdb header -- no channel creation records")
        return {}
    channel_names, handle_base = resolved
    with open(path, "rb") as f:
        fields = header_fields(f.read(128))
    data_slots, page_size = fields["data_slots"], fields["page_size"]
    symbols = read_blob_symbols(path)
    if symbols is None or not page_size:
        return {}
    candidates: Dict[int, List[ChannelMaker]] = {}
    with open(path, "rb") as f:
        for blob in iter_blobs(path):
            if blob.blob_index < data_slots:
                continue
            m = _CHANNEL_OBJECT_NAME_RE.fullmatch(symbols.get(blob.blob_index - data_slots, ""))
            if not m:
                continue
            channel_slot = int(m.group(1)) - handle_base
            if channel_slot not in channel_names:
                continue
            f.seek(blob.offset)
            data = f.read(min(blob.n_pages * page_size, 50_000_000))
            if len(data) < 28:
                continue
            end = 28 + struct.unpack_from("<i", data, 24)[0]
            maker = _decode_reg_maker(data[:end])
            if maker is not None:
                candidates.setdefault(channel_slot, []).append(maker)
    result: Dict[str, ChannelMaker] = {}
    for channel_slot, makers in candidates.items():
        name = channel_names[channel_slot]
        distinct = []
        for mk in makers:
            if mk not in distinct:
                distinct.append(mk)
        if len(distinct) > 1:
            _warn(
                f"{path}: channel {name!r} has {len(distinct)} different creation "
                f"records across stale/duplicate registry entries -- using the "
                f"last one in blob-chain order"
            )
        result[name] = makers[-1]
    return result

pygdb.ChannelMaker(tool, label, parameters) dataclass

How a channel was made: one MAKER record from the file's registry.

Attributes:

Name Type Description
tool str

The tool that created the channel, as the file records it -- a GX name ("newchan.gx", "gx\\copy.gx") or a .NET entry point ("geogxnet.dll(Geosoft.GX.MathExpressionBuilder...;RunChannel)").

label str

The tool's human-readable name ("New channel", "Copy channel").

parameters dict of {str : str}

The tool's parameters, {"TOOL.KEY": "value"} in file order -- e.g. {"NEWCHAN.DISPWIDTH": "10", ...} or the formula of a math expression. Empty when the tool recorded none.

Notes

[CONFIRMED] layout on all 305 corpus records (docs/spec.md section 9, docs/provenance/notes.md section 6.8c). Values are returned as the text the file stores; their meaning is the tool's, not decoded here.

pygdb.find_display_lists(path, channels=None)

Read the file's Display List objects.

Parameters:

Name Type Description Default
path str

Path to the .gdb file to scan.

required
channels iterable of ChannelRecord

The file's channel table, used to resolve each entry's handle to the channel's current name. If not given, this calls read_channels(path) itself.

None

Returns:

Type Description
list of list of DisplayListEntry

One list per live Display List object, in blob-symbol slot order, each in stored order. [] for a file with none, or with no readable blob-symbol table.

Warns:

Type Description
GDBParseWarning

If path doesn't start with the expected magic.

Notes

[CONFIRMED] layout on all 48 corpus instances (docs/spec.md section 9): a VV of fixed-width strings (82 or 130 bytes), each name\0handle\0. [LIKELY] meaning: the channels shown in the database's spreadsheet view. The live copy of each object comes from the blob directory (docs/spec.md section 2.2).

Source code in pygdb/registry.py
def find_display_lists(
    path: str,
    channels: Optional[Iterable[ChannelRecord]] = None,
) -> List[List[DisplayListEntry]]:
    """
    Read the file's `Display List` objects.

    Parameters
    ----------
    path : str
        Path to the `.gdb` file to scan.
    channels : iterable of ChannelRecord, optional
        The file's channel table, used to resolve each entry's handle to
        the channel's current name. If not given, this calls
        `read_channels(path)` itself.

    Returns
    -------
    list of list of DisplayListEntry
        One list per live `Display List` object, in blob-symbol slot order,
        each in stored order. `[]` for a file with none, or with no
        readable blob-symbol table.

    Warns
    -----
    GDBParseWarning
        If `path` doesn't start with the expected magic.

    Notes
    -----
    **[CONFIRMED]** layout on all 48 corpus instances (docs/spec.md
    section 9): a VV of fixed-width strings (82 or 130 bytes), each
    `name\\0handle\\0`. **[LIKELY]** meaning: the channels shown in the
    database's spreadsheet view. The live copy of each object comes from
    the blob directory (docs/spec.md section 2.2).
    """
    resolved = _channel_handles(path, channels)
    if resolved is None:
        _warn(f"{path}: not a readable .gdb header -- no display lists")
        return []
    channel_names, handle_base = resolved
    symbols = read_blob_symbols(path)
    if symbols is None:
        return []
    objects = _live_admin_objects(path)
    lists: List[List[DisplayListEntry]] = []
    for slot in sorted(k for k, name in symbols.items() if name == _DISPLAY_LIST_NAME):
        data = objects.get(slot)
        if data is None:
            continue
        at = data.find(_VV_TAG_BLOCK)
        if at == -1 or at + 20 > len(data):
            continue
        _zero, element_type, count = struct.unpack_from("<iii", data, at + 8)
        width = -element_type
        start = at + 20
        if width <= 0 or count < 0 or start + count * width > len(data):
            continue
        entries = []
        for i in range(count):
            parts = data[start + i * width:start + (i + 1) * width].split(b"\x00")
            label = parts[0].decode("latin-1")
            handle_text = parts[1].decode("latin-1") if len(parts) > 1 else ""
            if not handle_text.isdigit():
                continue
            handle = int(handle_text)
            entries.append(DisplayListEntry(
                label=label, handle=handle, channel=channel_names.get(handle - handle_base),
            ))
        lists.append(entries)
    return lists

pygdb.DisplayListEntry(label, handle, channel) dataclass

One entry of a Display List object.

Attributes:

Name Type Description
label str

The channel name stored in the list. A cached label: a channel renamed after it was added keeps its old name here.

handle int

The channel's global symbol handle, which identifies it.

channel str or None

The channel's current name, resolved from handle through the channel table; None if no current channel has that handle.

Warnings and errors

pygdb.GDBParseWarning

Bases: RuntimeWarning

Warned whenever this reader hits a blob, chunk, or record it can't parse.

Covers an unexpected byte sequence, a file that ends prematurely (truncated download, or a blob chain that runs past EOF), an administrative-blob variant it doesn't recognize, or anything else that doesn't fit the confirmed structure. This reader is designed to degrade gracefully rather than hard-crash on this whole class of problem: functions return whatever they successfully decoded up to the point of trouble (a shorter-than-expected list, an empty list, or in the worst case an empty result) instead of raising, and a GDBParseWarning describing what couldn't be decoded and why is always issued alongside, so a caller can tell a clean, complete result from a partial one and go investigate.

Notes

See docs/provenance/notes.md's "reader robustness" notes for the design rationale (an explicit engineering request, not new format research).

This does not apply to a handful of genuine precondition failures that aren't "this file has an interesting anomaly" (e.g. calling read_blob_values with a channel/blob pair that can't possibly match) -- those still raise normally.

pygdb.GDBUnseenFeatureWarning

Bases: UserWarning

Warned when a file has a format feature this reader has never seen.

Nothing is wrong with the file, and the data is read normally. The notice gives the exact values seen and asks the user to run unseen_feature_report and post its output on the project's issue tracker, which helps finish the format specification. Each feature is reported at most once per file per process.

Notes

Deliberately not a GDBParseWarning: that one means something could not be decoded. To turn these notices off, either filter the warning::

import warnings
import pygdb
warnings.simplefilter("ignore", pygdb.GDBUnseenFeatureWarning)

or set the environment variable PYGDB_UNSEEN_FEATURE_NOTICES=0, which also skips the checks that run when a file is opened.

pygdb.unseen_feature_report(path)

Describe every format feature in path that this reader has never seen.

This is what a GDBUnseenFeatureWarning asks you to run and post. It returns nothing about the data in your file -- only about how the file is laid out. But don't take our word for it: check the source code for yourself! It's this function and the _find_* checks above it, in pygdb/unseen.py.

Parameters:

Name Type Description Default
path str

Path to the .gdb file.

required

Returns:

Type Description
str

A plain-text report, ready to paste into the issue form. It never contains the file's path or name. With nothing unusual found, it says so.

Notes

Every check runs again, whatever notices were already shown, and no notice is issued while it does. It runs even when PYGDB_UNSEEN_FEATURE_NOTICES=0. What each feature adds beyond its one-line evidence:

  • header words: the 128-byte header (table sizes, counts, page numbers);
  • channel +108/+116: each flagged channel's 128-byte record (name, type, format, display and array widths);
  • user table: nothing, since user records hold user names and paths;
  • line type or line block: each flagged line's 128-byte record (name, category, date, number, flight, type, version);
  • projection method or slot 7: the projection record, +104..+652 (coordinate-system, datum, ellipsoid, transform, units and projection names; the method code; the parameters);
  • projection object member: each member's length;
  • MAKER field: the tool's name, never its parameters;
  • extension-object list: its payload, up to 512 bytes.

GDBParseWarnings raised while the file is read are left alone.

Examples:

>>> print(pygdb.unseen_feature_report("survey.gdb"))
Source code in pygdb/unseen.py
def unseen_feature_report(path: str) -> str:
    """
    Describe every format feature in `path` that this reader has never seen.

    This is what a `GDBUnseenFeatureWarning` asks you to run and post. It
    returns nothing about the data in your file -- only about how the
    file is laid out. But don't take our word for it: check the source
    code for yourself! It's this function and the `_find_*` checks above
    it, in `pygdb/unseen.py`.

    Parameters
    ----------
    path : str
        Path to the `.gdb` file.

    Returns
    -------
    str
        A plain-text report, ready to paste into the issue form. It never
        contains the file's path or name. With nothing unusual found, it
        says so.

    Notes
    -----
    Every check runs again, whatever notices were already shown, and no
    notice is issued while it does. It runs even when
    `PYGDB_UNSEEN_FEATURE_NOTICES=0`. What each feature adds beyond its
    one-line evidence:

    - header words: the 128-byte header (table sizes, counts, page
      numbers);
    - channel `+108`/`+116`: each flagged channel's 128-byte record
      (name, type, format, display and array widths);
    - user table: nothing, since user records hold user names and paths;
    - line type or line block: each flagged line's 128-byte record (name,
      category, date, number, flight, type, version);
    - projection method or slot 7: the projection record, `+104..+652`
      (coordinate-system, datum, ellipsoid, transform, units and
      projection names; the method code; the parameters);
    - projection object member: each member's length;
    - `MAKER` field: the tool's name, never its parameters;
    - extension-object list: its payload, up to 512 bytes.

    `GDBParseWarning`s raised while the file is read are left alone.

    Examples
    --------
    >>> print(pygdb.unseen_feature_report("survey.gdb"))  # doctest: +SKIP
    """
    from importlib.metadata import PackageNotFoundError, version

    from .gdb import GDB

    findings: List[Finding] = []
    token = _collector.set(findings)
    try:
        db = GDB(path)
        try:
            db.channels
            db.lines
        finally:
            db.close()
    finally:
        _collector.reset(token)

    try:
        pygdb_version = version("python-gdb")
    except PackageNotFoundError:  # pragma: no cover -- only when not installed
        pygdb_version = "unknown"
    lines = [
        "pygdb unseen-feature report",
        f"pygdb version: {pygdb_version}",
        "This report says nothing about the data in your file -- only about how",
        "the file is laid out: no survey values, user names, paths or processing",
        "parameters. But don't take our word for it: check the source code",
        f"yourself ({SOURCE_URL}),",
        "and read this before posting it.",
        "",
    ]
    if not findings:
        lines.append("No unseen features found in this file.")
    for finding in findings:
        lines.append(f"[{finding.feature}]")
        lines.append(f"  Evidence: {finding.evidence}")
        if finding.details:
            lines.append("  Details:")
            lines += ["    " + row for row in finding.details.splitlines()]
        lines.append("")
    return "\n".join(lines).rstrip() + "\n"

pygdb.GRDParseWarning

Bases: RuntimeWarning

Warned on a truncated file or a compressed block that fails to decompress.

See gdb_reader.GDBParseWarning for the same design rationale applied here: return whatever grid data was successfully decoded (padded with the header's own dummy value for the missing tail, so the array shape still matches shape_e * shape_v) rather than raising and discarding a whole grid over one bad/missing block.

pygdb.LZRW1DecodeError

Bases: Exception

Raised when a Speed-mode chunk can't be decoded.

Raised by decode_speed_chunk/parse_chunk_header for truncated/corrupt data, or a subtype/marker this module doesn't recognize. A single, deliberately narrow exception type (rather than a bare AssertionError/IndexError/struct.error grab-bag) so callers -- notably gdb_reader.read_blob_values -- can catch exactly this and fail gracefully (return whatever was already decoded elsewhere, emit a clear warning) instead of crashing.

Notes

See docs/provenance/notes.md's "reader robustness" notes for the design rationale; this reader is not meant to hard-crash on a truncated download or an unrecognized real-world variant, per an explicit engineering request from the coordinator.