pygdb¶
pygdb is a clean-room Python reader for Geosoft's proprietary .gdb
("Geosoft Database") binary format and the sibling .grd grid format.
No affiliation with Geosoft or Seequent
This project is an independent, third-party implementation. It has no affiliation with, and is not endorsed, sponsored, or certified by, Geosoft Inc., Seequent, Bentley Systems, or any of their successors. "Geosoft" and "Oasis montaj" are trademarks of their respective owners.
It was produced entirely by clean-room means: vendor-published
open-source code and documentation, independent third-party format
readers, one openly-specified independent successor format (.geoh5,
read via the third-party geoh5py library), and byte-level analysis of
real, publicly downloaded .gdb files. No Geosoft software of any kind
(the geosoft/gxapi/gxpy compiled package, Oasis montaj, Geosoft
Desktop, or the free Geosoft Viewer) was installed, imported, or
executed at any point in producing this library. See
Provenance for the full research trail.
Reading only โ writing or mutating .gdb/.grd files (Geosoft's own
proprietary formats) is out of scope. Exporting what's been read into a
different, openly-specified format is a separate concern and is
supported -- see Exporting to xarray and
Exporting to geoh5 below.
Installation¶
Package name
The distribution is named python-gdb on PyPI (pygdb itself was
already registered there for an unrelated project). The importable
package is still pygdb.
numpy is the one required dependency -- it's what lets read()
return correctly-shaped arrays rather than a flat, unshapeable buffer
(see Quick start below). Everything else (the Rust accelerator,
zensical for docs, pytest for tests) is optional.
Quick start¶
The GDB class is the recommended entry point: a name-based view over
a single .gdb file.
from pygdb import GDB
db = GDB("example.gdb")
db.compression # CompressionInfo(code=0, name='DB_COMP_NONE', ...)
db.coordinate_systems # ['NAD83 / UTM zone 11N', 'WGS 84'] (best-effort, may be [])
db.coordinate_channels # {'X': 'Easting', 'Y': 'Northing', 'Z': None} (from the
# file's own internal registry, may be all-None)
db.channel_settings # {'raw_mag': {'UNITS': 'nT'}, ...} (real per-channel
# registry values -- units, labels, formulas -- may be {})
db.projection_parameters # {'NAD83 / UTM zone 11N': ProjectionParameters(...)} (real
# datum/ellipsoid/projection numbers, keyed by coordinate_systems)
db.channel_makers # {'final_mag': ChannelMaker(tool='geogxnet.dll(...MathExpressionBuilder...)',
# parameters={...'CHANNELINPUTBOX': 'ch_13=ch_5 - ch_12;'...}), ...}
db.display_lists # [[DisplayListEntry(label='lat', handle=2070, channel='lat'), ...]]
db.line_names[:5] # ['L1000', 'L1001', 'L1010', 'L1020', 'L1030']
db.channels_on_line("L1000") # channels that actually have data on this line
db.read("L1000", "Easting") # random access by (line name, channel name)
# -> ndarray, shape (n_rows,) for a scalar
# channel, (n_rows, array_width) for a
# VA/array channel (docs/spec.md section 5)
for channel, values in db.iter_line("L1000"):
print(channel.name, values[:3])
The lower-level functions GDB is built on (read_channels,
read_lines, iter_blobs, find_blob, read_blob_values, ...) are
also exported directly from pygdb for anyone who wants slot-index-
based access or to walk the blob chain themselves.
See the format specification for the on-disk structure this library implements, with a confidence rating (confirmed / likely / guess / unknown) on every field.
Files with duplicate or unlisted blobs¶
The blob chain is append-only, so a .gdb can hold two blobs for the same
(line, channel): an older copy (in one real file, a re-sorted working copy)
can sit before or after the current one. The file records which blob is
current in its blob directory (see the specification),
and GDB reads that: an entry is used only if it points at a blob with the
right index and size. A duplicate the directory resolves is not reported.
- A blob the directory does not list as live is skipped, and one
GDBParseWarningsays how many and in which channels. In the project's test corpus that is two whole channels of one file. Why a blob is unlisted is not established, so to read them anyway:
- If a directory entry fails validation, or the file has no directory at all,
the last blob in chain order is used and a
GDBParseWarningsays so; that is a guess, and not always the current copy.
Unseen-feature notices¶
A few details of the format are still unknown (the specification's
ยง11), and most can only be settled
by a file that does something none of the project's sample files do: a
second database user, a projection method other than the four seen, a
channel scale factor other than 1. pygdb checks for these, and when a
file has one it issues a GDBUnseenFeatureWarning:
survey.gdb: this file has a feature pygdb hasn't seen before (projection
method code 7). The data was read normally. To help finish the format spec,
please run pygdb.unseen_feature_report(r"survey.gdb") (or: python -m
pygdb.report "survey.gdb") and paste its output into an issue at
https://github.com/jcapriot/pygdb/issues/new?template=unseen-feature.yml.
Please don't attach the file unless it's public.
Evidence: coordinate system '...', method code 7, slots [...]
Nothing is wrong with the file, and nothing is sent anywhere: the notice is a local warning, shown at most once per file and feature.
To report it, run the report and paste what it prints into the issue form:
The report lists every unseen feature in the file, each with its evidence and the surrounding structural bytes we need to decode it. It says nothing about the data in your file, only about how the file is laid out: no survey values, user names, paths or processing parameters, and not even the file's name. But don't take our word for it: check the source code for yourself, and read the report before posting it.
Most checks cost nothing. Opening a file also scans its administrative objects (projections, channel creation records), which adds a few tens of milliseconds on a large file. To turn the notices off, filter the warning, or set an environment variable, which also skips that scan:
Exporting to xarray¶
ds = db.to_xarray("L1000") # one line -> one xarray.Dataset
ds = db.to_xarray() # whole file (default) -> every line stacked
# along a new "line" dimension
ds["Easting"] # a plain (station,) DataArray -- (line, station) in
# whole-file mode
ds["ISPD"] # a VA/array channel -> (station, ISPD_bin), its
# own dimension, not shared with other array
# channels even at the same width (see below)
One Dataset per line by default, sharing a "station" dimension
across every channel; pass line=None (or call to_xarray() with no
argument) for the whole file instead, with a real "line" coordinate
(ds.sel(line="L1000"), ds.groupby("line")) and a secondary
"line_category" coordinate alongside it. A VA/array channel (see the
format specification's section 5) gets its own second
dimension named after the channel -- deliberately not shared with any
other array channel even when their widths happen to match, since
that's a coincidence, not a guarantee they mean the same thing. If two
channels share a name anywhere in the file's channel table (rare, but
structurally possible -- see channel()'s docs), the second one's
variable name is disambiguated as "name[1]" rather than silently
overwriting the first, with a warning explaining why -- resolved once
for the whole file, so the same channel's name never depends on which
line you're looking at.
In whole-file mode, every line has to share one common station count
and one common set of channels, unlike single-line mode's per-channel
dimension-splitting escape hatch -- a channel that decodes short, or
is simply absent on some lines (this format's normal sparse grid), is
filled with a value matching its own data type instead: NaN for
float channels, "" for string channels, and Geosoft's own published
per-type "no data" sentinel for every integer type (the same
convention real files already use for individual missing values, e.g.
rDUMMY=-1.0E32). Whichever value was used is recorded as that
variable's _FillValue attribute -- the standard CF/netCDF convention
name, so it's discoverable programmatically (including by
ds.to_netcdf(...)) rather than needing to be known in advance.
xarray is an optional dependency, imported only when to_xarray()
is actually called -- importing pygdb itself never needs it.
Exporting to geoh5¶
Unlike to_xarray (one line, in memory), this is whole-file and writes
directly to disk: one Points object per line (only for lines that
have real data), grouped under one ContainerGroup named after the
.gdb file. Each line's vertices come from an x_channel=/y_channel=
pair, left unset by default -- which first tries db.coordinate_channels
(the file's own internal registry of which real channel plays the X/Y/Z
role, confirmed present and correct on every one of this project's real
sample files), falling back to "Easting"/"Northing" only when that
registry doesn't confirm a role. Pass an explicit channel name to
override both. A line missing its resolved X or Y channel is skipped
entirely, with a warning, rather than guessed at. z_channel= works the
same way, except unresolved (no registry match, no fallback) just means
every vertex gets Z = 0.0 -- a missing elevation channel is normal
and never blocks export.
A VA/array channel is exported as one Data entry per column
("name[0]", "name[1]", ...), tied back together with a
PropertyGroup named after the channel -- .geoh5 has no Data type
that holds more than one value per vertex, so this mirrors the same
"one value per gate, grouped" pattern geoh5py's own built-in survey
types use internally for multi-gate data. Duplicate channel names are
disambiguated the same way to_xarray does ("name[1]", with a
warning), and a channel that decodes to a different row count than its
line's vertex count is skipped (that channel only), since .geoh5
Data must match its object's vertex count exactly -- there's no
per-channel-dimension escape hatch the way xarray has.
Coordinate-system/CRS export isn't attempted: db.coordinate_systems
only returns best-effort names, not real EPSG codes or projection
definitions, which isn't enough to populate .geoh5's CRS metadata
correctly. geoh5py is an optional dependency, imported only when
to_geoh5() is actually called.
Exporting to pandas¶
db.to_dataframe() # whole file -> one DataFrame, all lines
# concatenated, with "line"/"line_category"
# columns added to tell rows apart
db.to_dataframe("L1000") # one line only -> no "line" column;
# df.attrs has line_name/line_category/path
# instead, same three keys to_xarray uses
A VA/array channel is exported as one column per element
("name[0]", "name[1]", ...) -- the same flattening to_geoh5 uses,
for the same reason: there's no natural "one cell holds an array"
representation in a plain 2D table. Duplicate channel names are
disambiguated the same way ("name[1]", with a warning).
Row-count mismatches are handled differently here than in either
sibling export, deliberately: a channel that decodes shorter than its
line's other channels is padded with NaN out to match (with a
warning), not given its own dimension (to_xarray) or skipped
(to_geoh5) -- pandas has no structural reason to avoid this the way
the other two formats do, and NaN-padding a ragged column is completely
ordinary, idiomatic pandas. A channel present on some lines but not
others in whole-file mode needs no special handling at all: it's just
missing from that line's own frame, and pandas.concat's normal
union-of-columns behavior NaN-fills the gap automatically.
One real wrinkle worth knowing: a channel can happen to share a name
with a column this method always adds in whole-file mode ("line"/
"line_category") -- confirmed real, not hypothetical (a real sample
file has a channel literally named "line"). When that happens, the
channel's column is renamed "channel_<name>" instead, with a
warning -- "line" always means which survey line a row came from.
pandas is an optional dependency, imported only when to_dataframe()
is actually called.
Optional Rust-accelerated backend¶
The pure-Python code in pygdb/ is always the reference implementation
and always fully correct on its own. This package is built with
maturin so that an optional Rust extension (pygdb._native, source
under rust/ in the repository) rides along and is used automatically
when present -- it accelerates LZRW1 decompression and fixed-width
string decoding, the two real CPU-bound hot paths profiling found in
this reader, roughly 3-6x on real files, and decompresses DB_COMP_SIZE
(zlib) data ~11% faster than the stdlib fallback with one fewer copy,
using the flate2 crate on its zlib-rs backend. A platform/Python-version
combination with a published wheel gets it with a plain pip install
python-gdb, no extra step; elsewhere pip falls back to building from
source, which needs a Rust toolchain. For local development: pip
install -e ".[dev]" (from the repository root) compiles it as part of
the editable install. See rust/src/lib.rs for what's implemented, and
what was tried and benchmarked as not worth keeping.
Status¶
This is an early-stage (alpha) implementation, validated against 22
independent real .gdb files from 3 unrelated public-sector agencies
(USGS, the Geological Survey of Queensland, and the Ontario Geological
Survey) โ see the specification's validation scope and
the provenance notes for details. It has not
been tested against every real-world variant of the format, and the
specification will be updated as more example files are tested against
it. Bug reports with reproducible example files are very welcome โ see
Contributing.