The Geosoft .gdb binary format — reference specification¶
This is a reference document: it describes the on-disk structure of
Geosoft's .gdb ("Geosoft Database") binary format as currently
understood, organized by the file's actual pieces rather than by the
order they were discovered in. It is a distillation, not new research —
every claim here was already established in provenance/notes.md, and every
section below cites the provenance/notes.md section(s) with the fuller derivation,
byte-level evidence, and cross-validation. provenance/log.md has the full
chronological research trail (including dead ends) behind that.
This document was produced entirely by clean-room means: vendor-published
open-source code and documentation, independent third-party format
readers, one openly-specified independent successor format (.geoh5,
read via the third-party geoh5py library), and byte-level analysis of
real, publicly downloaded .gdb files. No Geosoft software of any kind
(geosoft/gxapi/gxpy compiled package, Oasis montaj, Geosoft Desktop,
or the free Geosoft Viewer) was installed, imported, or executed at any
point, in producing this document or anything it's based on.
Confidence markers (identical to provenance/notes.md — kept consistent
deliberately):
- [CONFIRMED] — verified against real file bytes, ideally
cross-checked against two or more independent files, or independently
corroborated by two unrelated public sources.
- [LIKELY] — a specific, falsifiable hypothesis that passed at least
one real test but wasn't independently cross-checked a second way.
- [GUESS] — a plausible pattern noticed in the data, not tested
against an independent prediction. Could be wrong.
- [UNKNOWN] — observed but not understood; raw facts recorded for
whoever continues this work.
Where a marker applies to only part of a claim (e.g. "confirmed in modern files, unknown in older ones"), that's spelled out rather than rounded up or down to a single label.
Validation scope, as of this writing: these findings have been
tested against 22 independent real .gdb files (2 USGS, 17 GSQ,
3 Ontario) from 3 unrelated agencies (USGS, the Geological Survey
of Queensland, and the Ontario Geological Survey), spanning roughly
1991–2020, 3+ airborne survey system vendors, file sizes from ~2MB to
~1.93GB, and all three DB_COMP_* modes — plus one independent, non-Geosoft cross-check via a
paired .geoh5 file read with the third-party geoh5py library. Full
provenance for every sample file is in provenance/notes.md §5. A full-corpus
sanity pass (provenance/notes.md §6.9) ran the complete reader — header,
symbol table, blob-chain walk, data decoding, VA/array channels, and
REG/IPJ registry scan — against all 22 files together in one run,
not just pairwise as each piece was developed: zero exceptions, and it
surfaced two genuine new findings (the first non-GS_DOUBLE array
channel, §5; REG/IPJ content isn't universal, §9) rather than only
confirming what was already known.
1. Conceptual model¶
A .gdb file is a database of channels (named, typed data columns —
e.g. Easting, raw_mag, LEI_Conductivity) recorded across multiple
lines (one flight line, ground traverse, or drillhole each). Every
line shares the same global channel namespace, but each line has its own
independent run of data for each channel it uses — a sparse 2D grid of
(line, channel) cells, most of which are populated for a real survey but
not all (channels can be scratch/abandoned, or only recorded on some
lines). Each channel's per-line data is indexed by a fiducial — for
scalar channels, one value per fiducial; for VA/array channels
(§5), a fixed-size vector of values per fiducial (e.g. a 24-gate decay
curve, or a 30-layer depth profile).
This conceptual model comes from vendor documentation (S5–S8 in
provenance/notes.md §1), not from byte analysis — but it's exactly the shape the
byte-level structure below turned out to have. provenance/notes.md §6.4
independently confirmed column-major storage: one channel's data
for one line sits in one contiguous run on disk, not interleaved
row-by-row with other channels.
2. File header¶
[CONFIRMED] magic; [LIKELY]/[UNKNOWN] for most individual
fields. Full derivation: provenance/notes.md §6.1.
All bytes below are little-endian; all offsets are absolute byte offsets from the start of the file.
| Offset | Type | Field | Status | Notes |
|---|---|---|---|---|
| 0–3 | 4 bytes | Magic, literal ASCII "!CBD" (21 43 42 44) |
[CONFIRMED] | Stable across every real file examined (23 files, 3 agencies, ~1991–2020). Meaning of the letters not documented anywhere found — possibly "Compressed Binary Database" or similar, unconfirmed. |
| 4–15 | 12 bytes | Fixed sub-block, 00 00 00 00 00 00 02 10 08 01 00 00 in the common case |
[LIKELY] format/version signature | Real exception, now common: bytes 4–7 (header word 4) read f0 f0 f0 f0 instead of zero on 13 of 45 public files: DB_Mag_Elaine_1003.gdb, East_Isa_VTEM_Inversion.gdb, and all 11 ground-gravity databases of the 2024 OpenEI BRIDGE Bell Flat delivery (whose aeromagnetic database reads zero). (Earlier text said bytes 8–11; the variant bytes are at 4–7.) It does not follow compression mode, table capacities, line category, or vintage (1991-2024). Against its same-delivery sibling DB_Mag_MountGordon_1003.gdb, the Elaine header differs only in this word and the page count (word 112). [UNKNOWN] what it means. |
| 24 | int32 | chans_max — channel-table capacity |
[CONFIRMED] | Proven by the SUPER-anchor structural test (§3.1 below / provenance/notes.md §6.2), not just by matching a documented default. |
| 28 | int32 | blobs_max — blob-symbol-table capacity |
[CONFIRMED] | Equals word 84 in 23 of 23 files (provenance/notes.md §6.1b). |
| 32 | int32 | cache — number of slots at the end of the blob directory (§2.2), which hold the free list of superseded blobs |
[LIKELY] | 100 in most files, matching the vendor's documented GXDB default (cache=100); larger in others (500, 1000, 2500, 3750, 5000, 10000). The directory array is exactly word 44 = word 60 + cache slots of 6 bytes. |
| 36 | int32 | lines_max — line-table capacity |
[CONFIRMED] | Equals word 88 in 23 of 23 files, and word 48 == lines_max × chans_max in 23 of 23. |
| 40 | int32 | users_max — user-table capacity |
[CONFIRMED] | Equals word 96 in 23 of 23 files. |
| 100 | int32 | page_size — the paging stride used elsewhere in the file (§5, §6) |
[LIKELY], but strongly corroborated | Matches the documented normal value (1024) in most real files; the compressed files seen use larger values (e.g. 32768), which independently turned out to be the real on-disk paging stride for those files (§6) — strong indirect confirmation. |
| 104 | int32 | index_size — the size in bytes of everything up to the end of the symbol tables (vendor DB_INFO_INDEX_SIZE) |
[CONFIRMED] | Exactly channel_table_start + (chans_max + users_max) × 128 + 8 in 23 of 23 files; the blob region then starts at the next page boundary (offset 108). (It was earlier read as close to but not exactly on the end of the symbol tables: the missing piece is the 8 trailing bytes.) |
| 108 | int32 | blob_start_page — the page number where the blob/data region begins |
[CONFIRMED] | Multiply by page_size (offset 100) to get the absolute byte offset of the very first blob header. Verified exactly on 16+ real files across every compression mode (§6.3). |
| 112 | int32 | Number of pages in the blob region | [CONFIRMED] | Equals file_size / page_size − blob_start_page in 23 of 23 files (also the sum of every blob's n_pages, where checked). |
| 116 | int32 | Lost pages: the page count of orphaned blobs (referenced by no data or registry slot) that the free list (§2.2) could not hold | [CONFIRMED] arithmetic | Exact on every file examined except one resized database (§2.2), 48 of 49 (the 23 below plus 26 added in Session 12, including two more non-zero cases, each predicted from the free list: long_valley_ed.gdb 1,148 and Brunt_mag_2017.gdb 354): 0 wherever the free list holds every orphan; 26,966 and 26 in the two files whose free list is full, equal to the pages of the orphans left out. Matches the vendor's name DB_INFO_LOST_SIZE. (A first test compared it with all orphaned pages, including listed ones, and wrongly ruled it out.) |
| 120 | int32 | comp_level — compression mode: 0=DB_COMP_NONE, 1=DB_COMP_SPEED, 2=DB_COMP_SIZE |
[CONFIRMED] | All three values directly observed in real files; see §7 for what each actually means on disk (and its real, honestly-documented exceptions). |
| 84, 88, 92, 96 | int32 (each) | Capacities of the blob, line, channel and user symbol tables, in the vendor's DB_SYMB_* order (BLOB=0, LINE=1, CHAN=2, USER=3) |
[CONFIRMED] | 92 == chans_max, 96 == users_max, 88 == lines_max, 84 == blobs_max in 23 of 23 files. |
| 72, 76, 80, 64 | int32 (each) | Running totals of those capacities: blobs, + lines, + channels, + users (= total symbol slots) | [CONFIRMED] | Exact cumulative sums in 23 of 23 files. |
| 44, 48, 52, 56, 60 | int32 (each) | Partition of the blob_index space (§6.1) and of the blob directory (§2.2): 48 = lines_max × chans_max (first index past the (line, channel) data blobs); 52 = 48 + blobs_max; 56 = 52 + users_max; 60 = 56; 44 = 60 + cache (the total number of directory slots) |
[CONFIRMED] arithmetic and directory slot count | Exact in 23 of 23 files. It explains why administrative/registry blobs are addressed at blob_index = lines_max × chans_max + slot: one slot per blob symbol. |
| 8–20, 68 | int32 (each) | Constant in every file examined: word 8 0x10020000 and word 12 264 (bytes 8–15 of the fixed sub-block above), words 16, 20 and 68 zero -- 49 of 49 files |
[UNKNOWN] meaning | Not resolved. The reader issues an unseen-feature notice for any other value. |
2.1 Layout of everything before the first blob¶
[CONFIRMED] for the blob directory, the line, channel and user tables and
the end marker; [LIKELY] for the position of the blob-symbol table.
Full derivation: provenance/notes.md §6.1b and §6.1c.
0 256-byte header (this section)
256 24 bytes, zero in every file examined
280 blob directory: word 44 slots × 6 bytes (§2.2)
b0 = l0 − blobs_max × 128 blob-symbol table: blobs_max × 128 bytes
(its first 32 bytes coincide with the last five
directory slots -- see §2.2)
l0 = c0 − 24 − lines_max × 128 line table: lines_max × 128 bytes
l0 + lines_max × 128 24 bytes (unexplained)
c0 channel table: chans_max × 128 bytes (§3.1)
c0 + chans_max × 128 user table: users_max × 128 bytes
… + users_max × 128 8 bytes; the end of this is word 104
zero padding to a page boundary; the first blob (word 108)
The line table has an exact position. c0 − 24 − lines_max × 128
reproduces, in all 23 real files, the same lines in the same order as
searching for it heuristically (find_line_table); in four files
(DB_Mag_1027, DB_Rad_1027, DB_Mag_1141, DB_Rad_1141) the slot numbers
differ by a constant −1, which is the indexing quirk of §3.2. pygdb.read_lines
uses this exact position, so its slot numbers are true and the blob-chain
calibration in pygdb.GDB is only a fallback for a file where the arithmetic
cannot be validated.
The real record boundaries: every symbol record begins with its name --
[CONFIRMED] (provenance/notes.md §6.2d). The offsets above, and those in
§3, measure each record from a point before its name: 8 bytes before for
channels and users, 32 bytes before for lines and blob symbols. Measured
from the name instead, the four tables are one contiguous run of 128-byte
records with no gaps, on 22 of 22 files:
280 + word44 × 6 blob symbols blobs_max × 128
c0 + 8 − lines_max × 128 lines lines_max × 128
c0 + 8 channels chans_max × 128
c0 + 8 + chans_max × 128 users users_max × 128
... ends exactly at word 104
The "24 bytes (unexplained)", the "8 bytes" before word 104 and the
blob-symbol/directory overlap above are all artefacts of the old
boundaries. A field the §3 tables place before the name belongs to the
previous record: line +0..+31 is the previous line's true +96..+127;
channel/user +0..+7 is the previous record's true +120..+127. The §3
tables keep the old offsets because pygdb.gdb_reader parses at them;
each gives the true offset where it matters.
The blob-symbol table names every administrative blob (§6.4/§9). Record
k, measured from its name:
| True offset | Field | Status |
|---|---|---|
+0 |
Name, NUL-terminated | [CONFIRMED] |
+76 |
Category: 0 (DB_CATEGORY_BLOB_NORMAL) when the symbol is live; bit 0x10000 set when the slot is free |
[CONFIRMED] 0 on exactly the 1,267 records that own a blob; [LIKELY] meaning of 0x10000 (all 1,050 name-bearing records with it own no blob) |
+84 |
The object's size, equal to the owning blob's own +24 field (§9) |
[CONFIRMED], 1,265 of 1,267 |
The owning blob is the one at blob_index = lines_max × chans_max + k, and
the directory (§2.2) gives its location. The names are:
- Fixed objects (22 of 22 files):
Line Selection,Display List,Database Extension Objectsand__dbreg. Rarer ones:__dbmeta,OE.DB_ACTIVITY_LOGandOE32.View. The name decides what the blob's+44holds ([CONFIRMED], every administrative blob in the corpus):
| Name | +44 |
Content |
|---|---|---|
__<n>, __dbreg |
REG\0 |
registry (§9) |
?\|IPJ_<X>:<Y> |
IPJ\0 |
projection (§8) |
Database Extension Objects |
EXT\0 |
an empty list -- one member (code LMSL) with no content, identical in every file (payload always 80 bytes, ending at +108). Bytes after +108 are leftovers |
__dbmeta |
META |
a typed metadata tree, zlib-compressed (§9) |
Line Selection |
(payload bytes) | one byte per line slot: lines_max rounded up to a multiple of 8 bytes of ff from +28 (that count is the blob's +24), in every file. There is no tag: the ff ff ff ff seen at +44 is line-selection bytes 16-19, and when the payload is shorter than 16 bytes (lines_max = 10, payload 16), +44 lies beyond it and holds leftovers. [LIKELY] a per-line selected flag, all lines selected |
Display List |
(the object frame's length) | a bare VV of fixed-width strings (§9): records of 82 or 130 bytes, each a channel name, NUL, the channel's handle as decimal text (lat\0 2070). [CONFIRMED] layout on every instance; the handle identifies the channel and the name is a cached label (renamed channels keep the old name). [LIKELY] the channels shown in the spreadsheet view |
OE.DB_ACTIVITY_LOG |
text | plain-text creation record: source path, Created: timestamp, Lines:/Channels: equal to the file's own lines_max/chans_max, Compression level: (4 Melinda Downs files, 2008). The text starts before +44, so the blob header overwrites its first bytes |
OE32.View |
text | plain-text [OASIS VIEW] configuration -- the old LINE "tag" false positive (§9) |
- Projections: ?|IPJ_<X>:<Y>, naming the coordinate-channel pair. |
||
- Per-symbol REG objects: __<n>, where n is a global symbol |
||
| handle. Words 72/76/80 (§2) are the running totals of the blob, line | ||
and channel tables, so handles [word 76, word 80) are channels (slot |
||
n − word 76) and [word 72, word 76) are lines. A line-handle |
||
| object is a per-line registry that always declares zero entries | ||
(744 of 744, preamble +124 = 0, §9). Its handle is a real line's slot |
||
| in every case: line slot 0 in most files, and 417 of 631 lines in | ||
Magnetic_Data.gdb. Any key-like bytes after its header are leftovers. |
||
| A "resized table" explanation (old channel handles shifted into the | ||
| line range) was tested and does not fit. |
2.2 The blob directory¶
[CONFIRMED] layout and entry form; [LIKELY] meaning of a zero entry
and of the cache slots. Full derivation: provenance/notes.md §6.1c.
The file records which blob is the current one for every blob index. It is
an array of word 44 slots of 6 bytes starting at offset 280, addressed by
blob_index (§6.1):
slot i at 280 + 6 × i : uint32 word, uint16 n_pages
live entry word = 0x80000000 | S, S = (blob's file offset − first blob's offset) / page_size
n_pages = the blob's own n_pages (§6.3)
empty cache word = 0x40000000, n_pages = 0
absent word = 0, n_pages = 0
slots [0, word 48) (line, channel) data blobs: slot == line_slot × chans_max + channel_slot
slots [word 48, word 52) registry blob-symbol slots (§9)
slots [word 52, word 56) user slots: empty in every file examined
slots [word 56, word 60) empty gap (word 60 == word 56)
slots [word 60, word 44) `cache` slots: the free list -- superseded/freed blobs, same entry form
Evidence, on all 22 corpus files and the supplied file:
- Entries are exact handles. All 116,287 non-zero data-slot entries in the
23 files carry flag
0x8and land on a blob header whoseblob_indexis the slot and whosen_pagesequals the entry's count; none fails. The 1,302 non-zero registry-symbol slots do too (flag0x8in 853,0xCin 449). - Coverage. In 21 of 22 corpus files every real (line, channel) blob is
listed (100%). In the remaining one (
SAMAGEM_CDI) 5,168 of 5,750 (89.9%) are; all 582 unlisted blobs are two whole channels (291 lines each, no duplicates) that still have channel-table records. In the supplied file 104 of 110 are listed, the six unlisted being in two channels. - It says which copy is current when a (line, channel) has several blobs. In the corpus all 345 duplicated pairs (139 + 206, in two files) are listed at the last copy in chain order. In the supplied file the current copy is often an earlier one: of 27 listed duplicated pairs, 19 point at the first copy and 8 at the last. All 13 pairs that could be labelled independently (8 against spreadsheet exports, 5 by other means) agree with the directory.
- A zero slot means the blob is not live: it was freed. Every chain
blob that no data or registry slot references is either on the free list
below or counted in header word 116, in all 31 files. That includes
SAMAGEM_CDI's two whole channels,afgrav.gdb's 13 single-line channels and the supplied file's six. Why a whole channel's data was freed (deleted, rewritten) is not recorded. - The
cacheslots are a free list -- [CONFIRMED] identity, [LIKELY] purpose. Entries have the same form, with flag0x8or0xC. Within one file's free list the flag is nearly always uniform. An empty slot is0x40000000. The evidence: - No free-list entry points at a blob that a data or registry slot references (0 of 117,450 in the first 22 files).
- Wherever the free list has room, it holds every orphaned blob
(38 of 38 on
AG106386, 172 of 172 onDB_EM_293). - In the three files whose free list is full, the orphans left out total exactly header word 116's page count (§2).
So the list records superseded or freed blobs, whose space can be
reused. The name cache comes from the vendor's GXDB creation
parameter. The reader does not consult these slots.
- The 0x4 bit flips on every rewrite -- [CONFIRMED] pattern, [LIKELY]
meaning. Pair each live entry with the freed copy of the same blob
index on the free list. The two carry opposite 0x4 bits in every
one of 696 pairs:
- 404 pairs with a live 0xC and a freed 0x8;
- 292 with a live 0x8 and a freed 0xC (139 of them data blobs).
Objects with two freed copies (14) have one of each. It reads as a
write-generation parity bit. Three consequences:
- An entry never rewritten keeps 0x8, e.g. 665 live registry entries
with no freed copy.
- The 26 live 0xC entries in Magnetic_Data.gdb with no freed copy
are exactly the 26 pages its full free list lost (header word 116).
- Live data entries are nearly always 0x8. The one exception is in
USGS OFR 2011-1270 Kalay_nk.gdb: a live 0xC data entry points at
the earlier of two copies, and the later copy is on the free list.
- A resized database is the exception to the free-list bookkeeping --
[LIKELY]. OpenEI BRIDGE GP_Master_Gravity_11082023.gdb has
lines_max = 100 and chans_max = 200, so data slots end at 20,000.
Yet it holds a complete set of 45 administrative objects (blob class
100: 42 registries, Line Selection, Display List, EXT) at blob
indexes 10,000-10,044, where they would sit if lines_max had been
50. It also holds 27 blobs of a line slot that is no longer a line.
None of these 72 blobs is live or on the free list (which has room),
and header word 116 reads 20 against 117 unlisted pages. So a table
resize leaves its leftovers outside the free list and the lost-page
counter. Every other file in the corpus follows the rules above.
Reader behaviour (pygdb.read_blob_directory, pygdb.GDB): an entry is
valid if its live bit (0x80000000) is set -- top nibble 0x8 or 0xC
-- and:
- its start page (the low 30 bits) lands on a blob header in the chain;
- that blob's
blob_indexequals the slot; - its
n_pagesequals the entry's count.
A valid entry selects the blob. (Until Session 12 the reader accepted
0x8 only. For the Kalay_nk.gdb entry above it warned and fell back to
the later, freed copy. There the two copies decode to identical values,
but in general that would pick the stale one.)
A zero entry in a file that has a directory means "not live": the blob is
skipped (a warning says how many, and GDB(include_unlisted_blobs=True)
reads them anyway). A non-zero entry that is not valid falls back to the
last blob in chain order, with a warning. A file with no directory (every
data slot zero, or header words inconsistent with the layout) is served
from the chain as before, last copy winning, with a warning if a pair is
duplicated.
No overlap. An earlier version of this section said the last five directory slots overlap the start of the blob-symbol table. That was an artefact of the old record boundaries: the table starts exactly where the directory ends (§2.1).
Not decoded: the purpose of the 0x4 rewrite bit, and the per-record
fields listed in provenance/notes.md §6.1c "What remains unexplained".
Not part of the header proper, but adjacent territory: REG/ coordinate-system (map projection) metadata — see §8 for what's now known.
3. The symbol table¶
[CONFIRMED] as the strongest single structural result on .gdb
itself. Full derivation: provenance/notes.md §6.2, §6.2b, §6.3.
3.0 Shared structure¶
Channel records, line records, and user records all share one
mechanism: a sequence of fixed 128-byte records — [CONFIRMED],
the same stride confirmed for all three record kinds. The vendor's own
published enum (DB_SYMB_BLOB=0, DB_SYMB_LINE=1, DB_SYMB_CHAN=2,
DB_SYMB_USER=3, provenance/notes.md §2) is consistent with this being one
unified symbol-table scheme with (at least) four symbol kinds, of
which only the LINE, CHAN, and USER regions have been located and
decoded — the BLOB (0) kind has only been observed as an unrelated
embedded projection dictionary at a different, 4096-byte stride
(provenance/notes.md §2.3 in provenance/log.md); it is not the same thing as the data
"blobs" in §6 of this document, despite the name collision (Geosoft's
own terminology overloads "blob" for both).
[CONFIRMED]: the channel table is immediately followed by the user
table (chans_max 128-byte channel records, then users_max 128-byte
user records) — this exact adjacency was the single cleanest structural
proof in the whole project (provenance/notes.md §6.2): the default super-user
name ("SUPER" or, in some real files, lowercase "super" — both seen)
sits at channel_table_start + chans_max × 128, and walking backward
from it by chans_max × 128 bytes lands exactly on the file's real
first channel.
3.1 Channel record layout (128 bytes)¶
| Rel. offset | Type | Field | Status |
|---|---|---|---|
+0..+7 |
8 bytes | Always zero (602 of 602 real channels). Not part of this record: it is the previous record's true +120..+127 (§2.1) |
[CONFIRMED] zero; [UNKNOWN] meaning |
+8 |
64 bytes | NUL-padded channel name (budget matches vendor's DB_SYMB_NAME_SIZE=64) |
[CONFIRMED] |
+84 |
int16 | Data-type code: positive = GS_* type (§4); negative = string byte-width (literal, not ×4) |
[CONFIRMED] |
+86 |
int16 | Matches the vendor's DB_ARRAY_BASETYPE_* enum values, but not reliably an array indicator on its own |
[LIKELY] name match, [UNKNOWN] exact write-time semantics — see §5 |
+92 |
int16 | Display format code (§4) — matches DB_CHAN_FORMAT_DATE/TIME exactly on real date/time channels, 0 (NORMAL) elsewhere |
[CONFIRMED] |
+94 |
int16 | Display width (vendor get_chan_width). Equals NEWCHAN.DISPWIDTH in the channel's own MAKER record (§9) on 32 of 32 channels created by newchan.gx, and the ASEG-GDF2 .dfn field width on 34 of 35 scalar channels of AG106386 |
[CONFIRMED] |
+96 |
int32 | Display decimals (vendor get_chan_decimal). Equals NEWCHAN.DISPDIG in the channel's MAKER record on 32 of 32, and the .dfn decimal count on 37 of 37 channels of AG106386 |
[CONFIRMED] |
+108 |
float64 | Exactly 1.0 on every genuine channel (1,247 of 1,247 in 49 files) |
[UNKNOWN] — plausible scale-factor field, never seen a non-1.0 value on a real channel |
+116 |
int16 | Exactly 5 on every genuine channel (1,247 of 1,247 in 49 files), regardless of the channel's own dtype at +84 — confirmed independent, not a copy of the type code, by checking it against every non-GS_DOUBLE dtype in the corpus (GS_USHORT incl. 512-wide array channels, GS_SHORT, GS_LONG, GS_FLOAT, and 9 string widths — all still read 5) |
[UNKNOWN] |
+118 |
int16 | Array width: number of elements per fiducial. 1 = scalar (the overwhelming majority); >1 = true VA/array channel |
[CONFIRMED] — see §5 |
Unused channel-table capacity (slots beyond the real channel count, up
to chans_max) is [CONFIRMED, with a documented revision] to be
cleanly zeroed in some real files (the 2020 USGS samples) but to hold
genuine leftover/uninitialized binary garbage in others (several real
1990s GSQ files) — a real reader must sanity-check candidate records
(NUL-terminated printable name; dtype/format codes in known valid
ranges) rather than assume clean padding. See provenance/notes.md §6.2 for the
real false-positive case this guards against. Array width (+118) must
also be at least 1. Three real GSQ files hold 16 leftover records with
clean names -- second copies of real channel names, projection-catalog
names -- and valid dtype and format codes, but width 0 and no data
(provenance/notes.md §6.2d). Leftover records also read +108 = 0.0
and +116 = 0.
3.2 Line record layout (128 bytes)¶
[CONFIRMED] existence, stride, and most fields now; a byte census
across all 22 real files (5,003 real line records, provenance/
notes.md §6.3b) resolved most of what was previously [UNKNOWN].
| Rel. offset | Type | Field | Status |
|---|---|---|---|
+0 |
int32 | Previous line's type — true +96 of the previous record (§2.1): DB_LINE_TYPE_*, 0 = NORMAL on every "L" line, 2 = TIE on every "T" line, 5,003 of 5,003 once read against the right line; 6 = RANDOM on the "D" line of afgrav.gdb |
[CONFIRMED] |
+4 |
int32 | Previous line's flight number (true +100) — mostly 0; where it varies, adjacent line pairs share values; no independent ground truth |
[LIKELY] |
+8 |
20 bytes | Previous line's true +104..+123. Always the identical byte pattern when populated (4,985 of 5,003 real lines) — a float32 -1e32 at +8, a float64 +1e32 at +20 (both the vendor's rDUMMY sentinels, §4), and a middle 8 bytes (+12) that decode exactly to the nearest float64 to a round -9×10^31 — not itself a catalogued vendor dummy. Read at the true offsets of every live line (5,575 lines in 49 files, normal and group), the pattern is identical on all of them. An earlier count through this previous-line view found 18 all-zero blocks, which were not traced to specific records |
[CONFIRMED] structure, real content never observed |
+28 |
int32 | Previous line's version (true +124) — the number after the dot in a repeat-line name: 1 for each of the 9 .1 lines, 0 for every other line, 5,003 of 5,003 |
[CONFIRMED] |
+32 |
up to 64 bytes | NUL-padded line name (note: not at +8 the way channel names are — line records reserve more leading fields) |
[CONFIRMED] |
+96 |
12 bytes | Always exactly zero, 5,003 of 5,003 | [CONFIRMED] reserved/unused |
+108 |
int32 | Category code — 100 matches DB_CATEGORY_LINE_NORMAL exactly on every real normal line seen; 200 (DB_CATEGORY_LINE_GROUP) also seen; a 65536 sentinel value seen on unused capacity slots |
[CONFIRMED] |
+116 (group lines) |
text | On a group line (category 200) the date and line-number positions instead hold a NUL-terminated group class name -- DB_Table on all 110 group lines of six USGS files (OFR 2011-1270). Group lines have freeform names (1000_points, Line_130pk, -100) and type 0. The vendor's set_group_class sets "the Class name for a group line", and group lines of one class share a list of associated channels; that list is the ASSOCIATED.<class> registry key (§9) |
[CONFIRMED] |
+112 |
int32 | Always exactly zero, 5,003 of 5,003 | [CONFIRMED] reserved/unused |
+116 |
float64 | A per-file, near-constant decimal-year timestamp (present on 18 of 22 files; the 4 oldest, 1991 GSQ, files have none) — reads as when the database itself was created/saved, not a per-line flight date (real lines in one file all share the identical value); the exact same value, byte for byte, as the user record's own +72 (§3.3) |
[LIKELY] |
+124 |
int32 | The line's own numeric line number — 4,994 of 4,994 real lines whose name ends in an integer match exactly ("L1150" → 1150); for a .1/.2 repeat-line name, holds the integer part only. The fraction is the line's version, stored in the next slot's +28 (its own true +124) |
[CONFIRMED] |
Measured from the name (§2.1), one line's own values are: name +0,
category +76, date +84, number +92, type +96, flight +100, the
20-byte dummy block +104, version +124. The number, version, type,
flight and date are exactly the vendor's DB_LINE_LABEL_FORMAT_* label
parts. The last line's +96..+127 falls on the old 24-byte gap plus the
first channel's old +0..+7. The first slot's old +0..+31 is the tail
of the blob-symbol table's last record.
Line-table physical slot numbering is 0-based, exactly like the channel
table, and — critically — this slot number is exactly the
line_slot_index used in the blob-addressing formula in §6.2.
[CONFIRMED] directly: physical slot 0 of the line table holds a
real survey's actual first line name on every file checked in the
original investigation (provenance/notes.md §6.6/§6.6d) — but not
universally, see the correction immediately below.
Correction, found while building a name-based reader on top of this
table (pygdb.GDB, §12): on a real GSQ file (rm001141), physical
slot 0 is a genuine, named record — "L0" — whose category code is
65636, not 100/200/65536. [LIKELY]: 0x10000 | 100, a
freed NORMAL line. The 0x10000 bit marks a free slot in the
blob-symbol table too, where it is set on exactly the records that own
no blob (§2.1). This slot has no data blob for any channel — it's a real
table entry, but not a usable survey line — and a scanner that only
recognizes categories 100/200 (as this specification's own
reference reader originally did) skips it, landing one slot late
and silently misnumbering every subsequent line for that file (a real,
found-by-testing bug, not a hypothetical one). The line table's exact
position (§2.1) avoids the problem altogether: the slot numbers come from
the table's true start, so the unrecognized slot 0 is simply skipped
without shifting the rest. Before that was known, pygdb.GDB corrected
for it by cross-checking candidate line numbering against which slots
actually have real blob data on disk (see its _calibrate_line_indices),
and still does when the exact position cannot be validated. The two
methods give identical line names and indices on all 22 corpus files.
3.3 User record layout (128 bytes)¶
Only one real user has ever been found in this project's entire
corpus: slot 0, the default superuser — 22 of 22 real files,
provenance/notes.md §6.2c. Records that first looked like more real
users in 3 files turned out to be leftover embedded-blob bytes landing
in unused table capacity by chance, indistinguishable from a name
until checked against the same clean-name test channel records already
needed (§3.1).
The default super-user's name sits at the same +8 offset as a channel
name (case varies by file: uppercase "SUPER" or lowercase "super"
both seen in real files — [CONFIRMED] both are the same structure,
not a format difference). The name shares its bytes with a path
string -- [CONFIRMED] on 11 of 22 files. A UTF-16LE file path starts
at +8 and runs for at most 32 characters (+8..+71, up to the
timestamp at +72). The ASCII user name and its NUL were written over
the path's first bytes: super\0 hides 3 characters, and everything
after them is intact. So the earlier reading of a 32-byte name field
ending at +40 was an artefact.
| Rel. offset | Type | Field | Status |
|---|---|---|---|
+8 |
NUL-terminated ASCII | User name, written over the start of the path below | [CONFIRMED] |
+8 |
up to 32 UTF-16LE chars | A path ending in the file's own name, e.g. …lder\1212\DB_AGG_1212.gdb, …\DB_EM_MountGordon_1003.gdb, …data\MLGRAV.gdb. Truncation rule: shorter than 32 characters, it is NUL-terminated (7 files, 17-31 characters); longer, it is cut at 32 characters (+71) with no terminator (4 files, e.g. …\DB_AGG_1213.gd). The visible strings begin mid-path, so the 3 hidden characters may be a ... ellipsis from path shortening -- [GUESS]. 11 files hold no such path |
[CONFIRMED] layout and truncation rule, on 11 of 22 files |
+72 |
float64 | The same per-file decimal-year timestamp as the line record's +116 (§3.2) — confirmed byte-for-byte identical on 2 real files checked directly |
[LIKELY] |
+84 |
int32 | Category. Empty and leftover user slots carry the free bit 0x10000 (sometimes with leftover low bits and leftover names such as SPF_250), like the other symbol tables, on every new file checked. Always exactly 131072 (0x20000) on every real superuser record. Measured from the name this is true +76, the category position in every symbol table (§2.1); DB_CATEGORY_USER_NORMAL is 0, so the 0x20000 bit is unexplained |
[CONFIRMED] value, [UNKNOWN] meaning |
+124 |
int32 | Always exactly -1 on every real superuser record |
[CONFIRMED] value, meaning open |
| everything else | — | Small varying integers and pointer-shaped values, no pattern found | [UNKNOWN] |
4. Data types, formats, and dummy values¶
Directly from vendor-published source (provenance/notes.md §2, source S3) —
[CONFIRMED] as literal values by definition, and independently
[CONFIRMED] to appear verbatim as real on-disk type/format codes.
| Code | GS_* type |
Width (bytes) | struct format |
|---|---|---|---|
| 0 | GS_BYTE (signed) |
1 | b |
| 1 | GS_USHORT |
2 | H |
| 2 | GS_SHORT |
2 | h |
| 3 | GS_LONG |
4 | i |
| 4 | GS_FLOAT |
4 | f |
| 5 | GS_DOUBLE |
8 | d |
| 6 | GS_UBYTE |
1 | B |
| 7 | GS_ULONG |
4 | I |
| 8 | GS_LONG64 |
8 | q |
| 9 | GS_ULONG64 |
8 | Q |
| 10–13 | GS_FLOAT3D/GS_DOUBLE3D/GS_FLOAT2D/GS_DOUBLE2D |
varies | not implemented in the reference reader |
A negative type code in a channel record (§3.1, offset +84) means
"string, -code bytes wide" — [CONFIRMED] directly against real
string lengths (a date channel formatted "2020/01/15", exactly 10
characters, has type code exactly -10). This is the literal on-disk
convention, and is simpler than the Python-layer convention in
vendor source (gx_dtype(), which computes -length*4 for a
different, UTF-8-safe allocation purpose) — where the two disagreed,
the real bytes settled it.
Display format codes (channel record offset +92):
| Code | Name |
|---|---|
| 0 | NORMAL |
| 1 | EXP |
| 2 | TIME |
| 3 | DATE |
| 4 | GEOGR |
| 5 | SIGDIG |
| 6 | HEX |
Dummy/no-data sentinel values (vendor-published, provenance/notes.md §2),
keyed identically to gdb_reader.GS_TYPE_NUMPY_DTYPE/GS_TYPE_DUMMY_VALUE:
| Type | Dummy value | Confidence |
|---|---|---|
iDUMMY (int32) |
-2147483647 |
[CONFIRMED] — appears verbatim in real decoded data |
rDUMMY (float32/float64) |
-1.0E32 |
[CONFIRMED] |
| signed byte | -127 |
[CONFIRMED] |
| unsigned byte | 255 |
[CONFIRMED] |
| signed short | -32767 |
[CONFIRMED] |
| unsigned short | 65535 |
[CONFIRMED] |
unsigned long (GS_ULONG) |
4294967295 (0xFFFFFFFF) |
[LIKELY] — vendor-published, matches the same enum's pattern exactly, but not yet independently observed as an in-file sentinel the way the others were |
signed 64-bit (GS_LONG64) |
-2**63 (0x8000000000000000) |
[LIKELY], same reasoning — .grd files apparently never use 8-byte elements in practice, so there's been no real data to check this against either |
unsigned 64-bit (GS_ULONG64) |
2**64 - 1 (0xFFFFFFFFFFFFFFFF) |
[LIKELY], same reasoning |
5. VA / array channels¶
[CONFIRMED], independently cross-validated against a public,
non-Geosoft standard. provenance/notes.md §6.2b.
A channel can store a fixed-size vector of values per fiducial
rather than a single scalar. This is marked by the int16 field at
channel-record relative offset +118 (§3.1): 1 = scalar (the
overwhelming majority of real channels), >N = an array of N
elements per fiducial (e.g. 24 for a multi-gate TEM decay curve,
30 for a layered-earth depth/conductivity profile, 50 for a
resistivity-inversion profile — all seen in real files).
Real survey data shows both possible representations of
conceptually identical "many values per station" data exist in the
wild: some processing pipelines flatten it into N separate scalar
channels (GEOTEMCh1..16, etc. — no array channels at all); others use
a genuine array channel. Both are real, valid, and independently
confirmed.
Independent cross-validation: the ASEG-GDF2 public ASCII standard
(a 2003 Australian industry format, completely unrelated to Geosoft)
uses an explicit Fortran-style repeat-count syntax (nFw.d) for array
fields. A real .dfn sidecar for a survey also delivered as .gdb
declares its array fields with the exact same repeat counts (30) that
the binary .gdb's +118 field reports for the same channel names —
agreement between two totally independent formats and toolchains.
A second field, relative +86 in the channel record, has values
matching the vendor's DB_ARRAY_BASETYPE_* enum (TIME_WINDOWS=1,
TIMES=2, etc.) but is not reliable as an array indicator by
itself — real arrays have been seen with +86=0, and real scalar
channels have been seen with +86=2 on every channel in a file
including obviously-scalar ones. [LIKELY] name match only;
[UNKNOWN] exact write-time semantics.
Non-GS_DOUBLE/GS_FLOAT array channels — [CONFIRMED] real.
Radiometric_Data.gdb (USGS) has two GS_USHORT array channels,
ISPD/ISPU, array_width=512 — a full airborne gamma-ray energy
spectrum recorded per station. Decoded values are physically correct:
zero counts in the lowest channels (below the detector threshold),
rising to a peak, then a smooth realistic decay — and row_count
(145,408) is exactly 284 stations × 512 channels. This channel pair
sat unnoticed (logged only as an ordinary scalar GS_USHORT channel)
from Session 1 until a full-corpus sanity pass re-decoded every
channel of every file at once (provenance/notes.md §6.9). No layout difference
from the GS_DOUBLE/GS_FLOAT case was needed to decode it correctly
— same array_width field, same flattened row_count × array_width
storage. Open gap: a string array channel has still not been
found in any real sample.
Reader behavior: GDB.read()/iter_line() reshape an array
channel's flat row_count-element buffer into a proper (n_rows,
array_width) numpy array before returning it -- the reader, not the
caller, is responsible for knowing array_width and reshaping
correctly (see gdb_reader._decode_numeric_or_string). A flat element
count that isn't a whole multiple of array_width (truncated/corrupt
data) warns and drops the incomplete trailing row rather than
returning a raggedly-shaped result.
Every array GDB.read()/iter_line() return -- numeric or string,
scalar or array-channel -- is a writable, independent ndarray, not a
read-only view: this project is fundamentally a file reader with no
inherent need to mutate decoded values itself, but a caller who wants
to is never blocked by an artificial restriction, and it's achieved
with no extra copy wherever that's actually possible (every case
except zlib-compressed numeric data, where the stdlib zlib module
has no API to decompress into a caller-supplied buffer, so one
explicit copy is paid there specifically to keep the result writable
too). String-typed channels (scalar or array) additionally decode to a
fixed-width Unicode dtype, <U{max_len}> (max_len = the longest
decoded record actually present, not the on-disk field width), not
dtype=object -- avoiding one Python str allocation per row, and
sized to the real content rather than a generously-oversized real
field (e.g. 64 bytes for a 5-character name) specifically because an
earlier version that used the on-disk width unconditionally measured
2.5x slower than the dtype=object approach it replaced. See
gdb_reader._decode_numeric_or_string's docstring for the full
reasoning and the exact numbers.
GDB.to_xarray(line) (optional xarray dependency) builds on this
directly: an array channel's array_width becomes a real, named
second dimension (f"{channel}_bin") on that channel's DataArray,
kept separate per channel even when two array channels happen to share
a width (real example: ISPD/ISPU are both 512-wide in
Radiometric_Data.gdb, but get ISPD_bin/ISPU_bin independently --
matching widths don't imply a shared semantic axis). See gdb.py's
to_xarray docstring for how it handles the same-line duplicate-
channel-name and mismatched-row-count edge cases this section already
documents as real, if rare, possibilities.
6. The blob index: locating (line, channel) → data¶
This is the format's core random-access mechanism, and the single
biggest structural question this project answered. [CONFIRMED]
end-to-end, for locating data in every compression mode.
provenance/notes.md §6.6/§6.6b/§6.6d.
6.1 The addressing formula¶
Every (line, channel) pair's data is stored in one blob — the
format's unit of per-line-per-channel storage. Each blob is identified
by a single integer, blob_index, computed directly from symbol-table
positions:
channel_slot_index— the 0-based physical slot number in the channel symbol table (§3.1).line_slot_index— the 0-based physical slot number in the line symbol table (§3.2).chans_max— the header's channel-table capacity (§2, offset 24).
[CONFIRMED] directly: verified against real ground truth on multiple channels (including a string channel) for the same real line, on 3 independent agencies' files, and structurally confirmed via a full-file scan on every real file tested.
6.2 The blob chain, and the directory of live blobs¶
Correction. An earlier version of this section said no table mapping
blob_index → file offset exists, having searched for one in the wrong place
(the blob-symbol table and a 12-byte reading of the cache). There is one: the
blob directory of §2.2 lists, for each blob_index, the current blob's
start page and size. It does not replace the chain (the directory is written
by the file, the chain is what carries the data and every stale copy of it),
but it is what decides which blob is current when a (line, channel) has
several (§11, issue #2).
Blobs are stored as a self-describing sequential
chain: each blob's own header records how many bytes it occupies, so
a reader locates the next blob purely by adding that size to the
current offset. To find a specific (line, channel) pair, walk the
chain from the start, comparing each blob's blob_index until it
matches (or build a full index once by recording every blob_index →
offset pair seen during one linear walk).
Where the chain starts: header offset 108 (§2) is a page number;
multiplying it by page_size (header offset 100) gives the exact byte
offset of the very first blob header. [CONFIRMED] on every real
file tested across all three compression modes.
Whole-file walk, exhaustively verified: starting from that offset
and repeatedly jumping forward by each blob's own declared size lands
exactly on the file's true byte size, with zero framing errors, on
every one of 20 real files tested (2MB to 1.93GB, all three
DB_COMP_* modes). This is the strongest form of confirmation this
project has produced for any single claim.
Pages that are not blobs -- [CONFIRMED], one real file. OpenEI BRIDGE
GP_Master_Gravity_11082023.gdb (2023) holds two kinds of such page:
- one page of leftover float data at the very start of the blob region;
- a run of 96 all-zero pages later on.
Blobs are contiguous around both, and walking on from the next page that
starts with the magic lands exactly on the end of the file. A reader must
therefore resynchronize at a page that is not a blob rather than stop
(pygdb.iter_blobs does, with one summary warning). The directory's start
pages count from the region start (header word 108), not from the first
blob found. This file was resized (§2.2), which is the likely origin of
the gaps. [LIKELY]
6.3 The plain blob header (DB_COMP_NONE, and "bare" blobs elsewhere — see §7.4)¶
48 bytes, always beginning with the same 4-byte magic. [CONFIRMED]
fields +0 through +12; [LIKELY]/[UNKNOWN] beyond that
(varies by file vintage).
| Rel. offset | Type | Field | Status |
|---|---|---|---|
+0 |
4 bytes | Magic CC CC 00 FF |
[CONFIRMED] |
+4 |
int32 | n_pages — this blob's total on-disk size, in pages (page_size) |
[CONFIRMED] — this is the authoritative field for chain-walking |
+8 |
int32 | A second value, usually equal to +4 |
[LIKELY] duplicate/allocated-vs-used field — not always equal to +4 (some real "administrative" blobs, §6.4, disagree); a correct reader must trust +4 alone, never require the two to match |
+12 |
int32 | blob_index (§6.1) |
[CONFIRMED] |
+16 |
int32 | Unix timestamp, or the sentinel 0x80000000 when unset. The sentinel is the norm in every vintage; timestamps appear on whole channels at once (Magnetic_Data.gdb: all 631 lines of 9 channels, 2020-03-25, the file's line date; one more channel the next day; derived channels unset) |
[LIKELY] when-written for imported channel data; not vintage-dependent |
+20 |
int32 | Blob class: 100 administrative, 200 plain data, 202 compressed data |
[CONFIRMED]: 202 on all 10,429 data blobs with the compressed-chunk magic after the header and 200 on all 106,681 without, both compression modes; 100 on every administrative blob, corpus-wide |
+24..31 |
8 bytes | Zero in some real files, non-zero in others | [UNKNOWN], varies |
+32 |
float64 | 1.0 in modern files examined (matches the channel-record +108 "scale factor" convention) |
[LIKELY] in modern files |
+40 |
int32 | Row count for this specific blob | [CONFIRMED] in modern files (decoding exactly this many values reproduces real, ground-truth-matching data); [UNKNOWN] in older files, where this offset decodes nonsensically |
+44 |
int32 | GS_* type code for this blob's data, matching the owning channel's own symbol-table dtype |
[CONFIRMED] in modern files; same caveat for older files |
+48 |
— | Real data begins here | [CONFIRMED] |
Real row data occupies row_count × element_width bytes starting at
+48; any remaining bytes out to n_pages × page_size are padding
(observed as zero).
6.4 The reserved/administrative blob variant¶
A real, recurring class of blob: blob_index decomposes to an
implausibly large "line number" (values in the hundreds to low
thousands seen, well past any real survey's line count), and offset
+44 (or the equivalent compressed-header field, §7.3) reads a
specific non-GS_* constant, 4670802, instead of a valid type code.
These are cleanly distinguishable and safely skipped by a reader
(negative or implausible row_count, or the tell-tale 4670802
constant).
The 4670802 constant is explained — [CONFIRMED]. It is not a
sentinel value: read as bytes rather than an int32, it is exactly the
ASCII string "REG\0". An administrative blob's own type-code field
holds the first 4 bytes of its own 3-letter object name instead of a
GS_* type code (an IPJ-tagged blob reads 49 50 4a 00 = "IPJ\0"
the same way) — see §9 for what that name introduces.
7. Compression¶
[CONFIRMED] for all three modes: which algorithm, exact on-disk
framing, and (for the common single- and multi-page cases) full
decoding verified against real ground truth. provenance/notes.md §3, §6.5,
§6.5b–f, §6.6b, §6.6d.
The header's comp_level field (§2, offset 120) declares one of:
| Value | Mode | Real on-disk algorithm |
|---|---|---|
| 0 | DB_COMP_NONE |
No compression — raw values directly after the plain 48-byte blob header (§6.3) |
| 1 | DB_COMP_SPEED |
LZRW1 (Ross Williams' 1991 algorithm) — not zlib, despite vendor documentation claiming otherwise for both tiers |
| 2 | DB_COMP_SIZE |
zlib/deflate |
7.1 The shared 16-byte "page primitive" magic¶
Both compressed modes — and the sibling .grd grid format — share one
low-level container primitive: a 16-byte sub-header immediately
preceding a compressed payload:
subtype is 1 for DB_COMP_SPEED payloads, 2 for DB_COMP_SIZE
payloads — [CONFIRMED] with zero exceptions across thousands of
real instances. reserved tracks the subtype exactly -- [CONFIRMED]:
0 with subtype 1 on all 7,015 LZRW1 chunk headers and 1 with subtype
2 on all 3,529 zlib ones, across .gdb and .grd files (the five
largest .gdb files were not scanned). The only other values seen come
from the 8-byte magic occurring by chance inside data. What it means
beyond that is unknown.
7.2 DB_COMP_SIZE framing (zlib)¶
Immediately after the 16-byte magic (§7.1), the zlib stream begins
directly — no further sub-header. [CONFIRMED] by decompressing
real streams with nothing but Python's standard-library zlib and
matching known ground truth exactly (a real constant column decoding
to 5027, matching an independently-sourced ASCII export digit for
digit).
7.3 DB_COMP_SPEED framing (LZRW1)¶
Immediately after the 16-byte magic, a further 12-byte length sub-header:
chunk_length includes these 12 bytes (chunk_length - 12 is the
number of raw bytes that follow). marker is a real flag, not just a
validation sentinel — [CONFIRMED], exactly two values seen across
all real files, zero exceptions:
marker value |
Meaning |
|---|---|
0xF4E5D6C7 (-186263865) |
Payload is genuine LZRW1-compressed data |
0xF0E1D2C3 (-253635901) |
Payload is stored raw, uncompressed (LZRW1 didn't shrink it, so the encoder gave up and stored it verbatim — Ross Williams' reference implementation's own FLAG_COPY case, re-purposed into this marker field) |
The compressed payload itself is [CONFIRMED] to be Ross Williams'
canonical LZRW1 algorithm, byte-for-byte — same 2-byte control word +
1-byte literal / 2-byte nibble-packed copy-item scheme as his own
public-domain reference implementation, with no 4-byte FLAG_BYTES
prefix (the reference C wrapper's convention; not carried into
Geosoft's on-disk format — the equivalent signal lives in the marker
field above instead). Validated exhaustively (every chunk, not a
sample) against all real Speed-mode files with the chunked scheme:
6,995 chunks, zero failures.
A blob is a chain of chunks — [CONFIRMED]. The 16-byte magic plus
12-byte sub-header above describe the first chunk only. A chunk
decompresses to at most 16368 bytes (2046 float64 values); a
channel holding more data than that on one line is split across
several chunks stored back to back. Every chunk after the first has
no magic of its own — just its bare 12-byte sub-header, immediately
followed by its payload, starting chunk_length bytes after the
previous chunk's sub-header began:
Each chunk is decoded independently (an LZRW1 back-reference never
reaches across a chunk boundary), and the outputs are concatenated. The
bytes after the last chunk are ordinary page padding and are not
zeros (non-zero on most real blobs checked), so they can't be used to
find the end of the chain — the blob header's total decompressed size
(§7.4, +24) is what tells a reader when to stop. Checked on every
real Speed blob in this project's corpus (7,015 blobs, 1,656 of them
multi-chunk) and on every real-line blob of a separately supplied, much
larger file (all of them multi-chunk): the chain's decompressed
lengths sum to exactly the header's +24 total, with zero exceptions.
A reader that decodes only the first chunk silently truncates every
channel longer than 2046 float64 values per line (2046 rows, or fewer
for wider element types) — the bug behind this correction.
Why some chunks are stored raw instead of compressed — [CONFIRMED]
to be a per-chunk data-compressibility outcome, not a size effect.
Tested and refuted a specific hypothesis (do smaller channels get
stored raw regardless of mode?) with a direct structural argument:
within one line, every channel shares the same row count, so two
same-size chunks of different channels can and do land on opposite
sides of the compressed/stored-raw split in the very same file. The
real driver is whether LZRW1 actually found exploitable redundancy in
that specific block's bytes (smooth, slowly-varying data like
projected coordinates compresses well; data with real low-order sensor
noise, like some raw sensor channels and derived correction channels,
often doesn't) — exactly matching Ross Williams' reference
FLAG_COMPRESS/FLAG_COPY design intent.
7.4 The blob header for compressed data¶
For a compressed blob, the header preceding the 16-byte page-primitive
magic (§7.1) is 56 bytes, not 48 (8 bytes more than the plain
DB_COMP_NONE header, §6.3). [LIKELY], checked by hand on real
records, not exhaustively decoded:
| Rel. offset | Field | Status |
|---|---|---|
+0..+15 |
Same magic/n_pages/n_pages_dup/blob_index layout as the plain header |
[CONFIRMED] |
+24 |
Total decompressed size of the blob, in bytes, across every chunk (§7.3). For a single-chunk blob this equals the chunk's own decompressed_length. DB_COMP_SIZE blobs too: their one zlib stream decompresses to exactly this many bytes |
[CONFIRMED] — 7,015 Speed blobs and 3,414 Size blobs in the corpus, plus every real-line blob of a separately supplied file, zero exceptions |
+28 |
16 + Σ chunk_length over the whole chain — the chain's total on-disk span including the first chunk's magic |
[CONFIRMED] — same 7,015 Speed blobs and the separately supplied file's, zero exceptions |
+40 |
float64 1.0 (same scale-factor convention as elsewhere) |
[LIKELY] |
+48 |
Real row count of the whole blob (+24 ÷ element width, for numeric types) |
[CONFIRMED] for numeric channels — 7,015 of 7,015 corpus Speed blobs |
+52 |
GS_* type code |
[LIKELY] |
+56 |
The 16-byte page-primitive magic (§7.1) begins here | [CONFIRMED] |
A real third on-disk blob variant — "bare" blobs. Some individual
blobs inside a genuinely-compressing file carry no chunk wrapper at
all: no 16-byte magic at the expected +56 position, just the plain
48-byte header (§6.3) with raw, uncompressed data straight after it —
indistinguishable in layout from a DB_COMP_NONE blob, just sitting
inside a file whose header declares real compression. [CONFIRMED]
real and correctly decodable this way. A correct reader must probe
for the 16-byte magic at the expected offset rather than trust the
file's declared comp_level for any individual blob.
7.5 Multi-page compressed blobs¶
[CONFIRMED]: a compressed blob spanning more than one
page_size-sized page is simply one continuous compressed stream
that spans across the page boundary — not one independently-framed
chunk per page. This was tested directly (checking whether page 2 of a
real multi-page blob starts with its own copy of the 16-byte magic —
it does not) rather than assumed from the single-page case. Reading
the entire n_pages × page_size span (minus the header) and handing
all of it to the decompressor in one call (zlib.decompressobj() for
DB_COMP_SIZE, correctly finding the real end of stream and reporting
the rest as harmless page padding; the LZRW1 chunk's own
decompressed_length/chunk_length fields for DB_COMP_SPEED,
already agnostic to page boundaries) decodes correctly — but note that
for DB_COMP_SPEED the span holds a chain of chunks, not one (§7.3),
so "one call" means walking the chain up to the header's total, not
decoding the first chunk and stopping. Verified on
real blobs up to 47 pages, both compression modes, including a
36-page array-channel blob whose decoded values matched independent
ground truth exactly.
7.6 Whole files/blobs that declare compression but contain none¶
[CONFIRMED] as a real, recurring phenomenon, [UNKNOWN] why.
Several real files declare comp_level=1 (DB_COMP_SPEED) but contain
zero compressed chunks anywhere — every blob is stored exactly like
a DB_COMP_NONE file. This is now confirmed on 4 real files spanning
an 80×+ size range (9.6MB to 807MB) and 2 unrelated deliveries,
directly ruling out file size as the explanation. The
comp_level header field reflects the mode the database was
configured with, not a guarantee that any particular byte was
actually compressed with it.
8. Coordinate-system (IPJ) metadata¶
[CONFIRMED] located, and — since this session's "REG " framing
work carried over directly — [CONFIRMED] for the fixed geodetic
parameter and name offsets too, on 3 independent agencies. provenance/
notes.md §6.7, §6.7b.
Per-database map-projection metadata is not a separate structure —
it lives inside the same "reserved/administrative blob" mechanism
described in §6.4, reached through the ordinary blob chain (§6.2) but
addressed with an out-of-range line_slot (values in the low
thousands — 1000–1002 and 2000 seen in real files) that acts as
a namespace for non-survey-data metadata rather than real per-line
data.
An IPJ blob shares the exact same 128-byte preamble as a REG blob
(§9) — the 0xff 0x00 0xe1 0x1e constant, the 0x00 0x1a 0xcc 0xff
separator, all at the identical offsets — confirmed on every IPJ
instance checked. Where REG's first nested tag is "REG ", IPJ's is
the already-known " JPI" name marker. (An earlier reading of a
"second nested tag" at +112, e.g. " UTM", "MGA ", is withdrawn:
+112 lies inside the name field below, and those were fragments of the
name text itself.)
The object's content past the 128-byte preamble is a fixed-offset
binary record, not the flat key/value form REG mostly uses. Real
geodetic parameters and names sit at the same absolute byte offset
(relative to the blob's own start) in every instance checked — first
confirmed on 30 of 30 real IPJ objects across 5 files/3 agencies, then
verified corpus-wide (63 of 63 real instances, all 22 real files):
| Offset | Field | Confidence |
|---|---|---|
+180 |
Datum name (NUL-terminated ASCII) — "GDA2020", "WGS 84", "NAD83", "NAD83(CSRS)" seen |
[CONFIRMED] |
+244 |
Ellipsoid name (NUL-terminated ASCII) — "GRS 1980", "WGS 84" seen |
[CONFIRMED] |
+308 |
Semi-major axis, float64 | [CONFIRMED] — 6378137.0 on every WGS 84 / GRS 1980 instance; 6378206.4 (Clarke 1866, NAD27, Alaska DGGS fortymile_linedata.gdb) and 6378388.0 (International 1924, Herat North datum, USGS GDR_clmag.gdb), each matching the published ellipsoid |
+316 |
Eccentricity, float64 | [CONFIRMED] — matches the ellipsoid at +244 exactly |
+324 |
Prime meridian, float64, degrees from Greenwich (the fourth value of a GXF datum string) | [CONFIRMED] position, 0.0 on 106 of 106 instances |
+332 |
Datum-transformation name (NUL-terminated ASCII) — "GDA94 to WGS 84 (1)", "NAD83 to WGS 84 (1)", "NAD83(CSRS98) to WGS 84 (1)" seen |
[CONFIRMED] on all 36 of 36 real instances corpus-wide that define one (a datum already stated in WGS 84 has nothing here to name); every one of the 3 agencies agrees |
+396..+451 |
The datum transformation's 7 Bursa-Wolf parameters, float64: dX, dY, dZ (metres), Rx, Ry, Rz (radians), scale (a multiplier, 1 + ppm/10⁶) |
[CONFIRMED] against the registry's _PJ_DATUM_TRANSFORM text on four datums: AGD66 to WGS 84 (12) reads −129.193, −41.212, 130.73 m, rotations converting exactly to 0.246/0.374/0.329 arc-seconds, and scale 0.999997045 (−2.955 ppm). The text is generated from these: it carries the float noise of the conversion (-2.95500000002669). With no transform (empty name field, 5 objects), dX is rDUMMY and the rest hold their defaults, 0 and scale 1 |
+452 |
Units name, NUL-terminated in a 64-byte field — m on every projected object (68), dega (degrees) on every geographic one (38) |
[CONFIRMED] |
+516 |
Units factor to metres, float64 (m,1 in GXF) |
[CONFIRMED] position, 1.0 everywhere |
+524 |
Projection name without the datum, NUL-terminated in a 64-byte field — UTM zone 11N, Australian Map Grid zone 54, *bas_polar; empty for geographic objects. The field ends exactly where the parameters begin |
[CONFIRMED], 106 of 106 |
+168 |
Projection method code, int32: 1 geographic (datum only), 11 Transverse Mercator, 3 Lambert Conic Conformal (2SP), 14 Polar Stereographic |
[CONFIRMED] for these four -- 1 and 11 throughout the corpus, 3 on two Lambert objects in two files, 14 on two objects in Brunt_mag_2017.gdb |
+588..+651 |
8 float64 parameter slots; their meaning depends on the method at +168 (table below) |
[CONFIRMED] layout |
Parameter slots by method -- [CONFIRMED]:
| Slot (offset) | Transverse Mercator (11) |
Lambert Conic Conformal 2SP (3) |
Polar Stereographic (14) |
|---|---|---|---|
0 (+588) |
latitude of natural origin | latitude of first standard parallel | latitude of natural origin |
1 (+596) |
longitude of natural origin | latitude of second standard parallel | longitude of natural origin |
2 (+604) |
unused | latitude of false origin | unused |
3 (+612) |
unused | longitude of false origin | unused |
4 (+620) |
scale factor at natural origin | unused | scale factor at natural origin |
5 (+628) |
false easting | easting at false origin | false easting |
6 (+636) |
false northing | northing at false origin | false northing |
7 (+644) |
unused | unused | unused |
Unused slots hold rDUMMY (-1.0e32, §4), and a datum-only object has all
eight unused. The slots carry the same values, in the same order, as the
registry's _PJ_PROJECTION text. Transverse Mercator's
"Transverse Mercator",34,66,0.9996,0,0 (afgrav.gdb, whose readme states
"Base latitude = 34 degrees N") fills slots 0, 1, 4, 5, 6.
"Lambert Conic Conformal (2SP)",30,38,0,66,0,0 fills slots 0, 1, 2, 3, 5, 6.
The names of the Lambert slots follow the EPSG parameter order for that
method (standard parallels, then false origin, then false easting and
northing), which fits the values (parallels 30° and 38° bracketing
Afghanistan, central meridian 66°E). A second, independent Lambert object
(British Antarctic Survey Brunt_mag_2017.gdb, *Weddel_lamb) reads
parallels −82/−78, origin −80, central meridian −81 in the same slots.
Polar Stereographic's text "Polar Stereographic",-71,0,0.994,0,2082760.109
fills slots 0, 1, 4, 5, 6 like Transverse Mercator, and that survey's own
metadata calls −71 the standard parallel.
The slot names are [CONFIRMED] by Geosoft's own GXF Revision 3
specification (Table 1, "Projection Transformation Methods"). It lists
each method's parameters "in the order required", taken from EPSG's
enumerated parameter order "with unused parameters omitted". The text
form omits unused parameters; the binary keeps them as unset slots. For
Polar Stereographic, GXF names slot 0 the latitude of natural origin, which
the British Antarctic Survey's own metadata calls the standard parallel.
Table 1 also lists methods not yet seen here (Hotine and Laborde Oblique
Mercator, Lambert Conic Conformal (1SP), Mercator (1SP) and (2SP), New
Zealand Map Grid, Oblique Stereographic, Swiss Oblique Cylindrical,
Transverse Mercator (South Oriented), *Albers Conic, *Equidistant
Conic, *Polyconic). Their parameter names are known, but not their
method codes or slot positions. On every Transverse Mercator object in the rest
of the corpus, slot 0 reads 0.
+588..+651 is one vector of 8 float64 projection parameters --
[CONFIRMED], ending exactly where the first IPJ member does (+652,
below). The registry's _PJ_PROJECTION text lists the vector's set
slots, in order -- [CONFIRMED], 43 of 43. Every corpus object with
matching text (matched by _PJ_NAME; 43 of 50 projected objects) has
exactly as many text values as set slots, equal in order: 41 Transverse
Mercator, 1 Lambert (2SP), 1 Polar Stereographic. So for a method whose
code has no known layout, its text still names its values: the text
gives the method name, Table 1 gives that method's parameter names in
order, and the binary's set slots give the values. Slot 7 is unset on
every object.
A clean, self-consistent confirmation, not a gap: an IPJ object
that defines only a datum/ellipsoid (no projection) reads the real
rDUMMY sentinel at +596 onward instead of a real number — the same
documented dummy-value convention used throughout this format (§4),
here correctly marking "not a projected system" rather than being
undecoded garbage.
Independent ground truth, exactly as before, now pinned to exact
offsets instead of "consecutive small deltas": in AG106386, the six
float64 values implied by the paired ASEG-GDF2 .prj sidecar's declared
projection match verbatim at +308/+316/+596/+620/+628/+636.
The same check against each file's own real, independently-known datum
holds on Ontario (MLMAG.gdb: NAD83/GRS 1980, central meridian
-81, false northing 0 — the northern-hemisphere UTM convention) and
USGS (Magnetic_Data.gdb: WGS 84, central meridian -117).
The name field and the type word -- [CONFIRMED] layout, 63 of 63.
+104: the coordinate-system name, NUL-terminated in a 64-byte field (+104..+167, the vendor'sDB_SYMB_NAME_SIZE). Every name in the corpus fits. The bytes after the NUL are uninitialized memory, which is where the "pointer-shaped" values once noted at+136..+176come from. They are not data.+168: the projection method code (table above).+172..+179is zero on every instance. The values do not match the vendor'sIPJ_TYPE_*constants (0-6).
Confirmed on 3 independent agencies, each geographically correct for its real survey location:
| File (agency) | Real projection name(s) found |
|---|---|
AG106386_Northern Georgetown_Conductivity.gdb (GSQ) |
"WGS 84 / UTM zone 54S" |
Magnetic_Data.gdb (USGS) |
"NAD83 / UTM zone 11N", "GRS 1980", "NAD83 to WGS 84 (1)" |
MLMAG.gdb (Ontario) |
"NAD83 / UTM zone 17N", "GRS 1980", "NAD83 to WGS 84 (1)" |
A confirmed micro-pattern for how a name is introduced: the 4-byte
tag " JPI" (a space plus what's plausibly the tail of the literal
string "IPJ" read across an alignment boundary) followed by an
int32 (1 in every instance seen) and then a NUL-terminated name
string. [CONFIRMED] directly on the working projected-CRS name in
every file checked.
Implemented as pygdb.registry.find_projection_parameters /
GDB.projection_parameters, returning a ProjectionParameters per
working coordinate-system name (the offset table above, mapping the
rDUMMY sentinel to None on the four projection fields rather than
returning it as a raw float). Parameters are read by projection method:
method_code is +168, and the named fields (latitude_of_origin,
central_meridian, scale_factor, standard_parallel_1/2,
false_easting, false_northing) come from that method's slots in the
table above. For a method code without a known layout, the named fields
stay None, and all eight raw slots are always available as
parameters. (Until Session 12 it read Transverse Mercator positions
for every object, and reported central_meridian=38 for the Lambert
system in afgrav.gdb.)
It also returns the method's GXF name (method) and its parameters keyed
by Table 1 names (method_parameters, e.g. latitude_of_natural_origin,
false_easting), with parameter_source saying where the names came
from:
"text": the file's own_PJ_PROJECTIONtext. It is used only when its values equal the binary's set slots in order, and, for a known method code, when it names that code's method. Otherwise the reader warns and ignores it. This is what decodes a method whose code has no known layout."binary": the confirmed slot layout for method codes 1, 3, 11 and 14.None: neither is available. Onlyparametersholds the values.
The values are always the binary's, never the text's. The text is
generated from the binary. Also returned: prime_meridian (+324), the
datum transform (+396) converted to GXF units (arc-seconds, ppm) as
datum_transform_parameters, units_name/units_factor (+452/+516)
and projection_name (+524). The older named fields are unchanged and
still filled only for the four known codes.
The IPJ object is a chain of member frames -- [CONFIRMED], 63 of 63. It uses the same framing as a registry (§9).
- The object frame (
ff 00 f0 0fat+28) holds 2 to 6 member frames (ff 00 e1 1e), starting at+60. - Each member's length counts from 28 bytes after its own length field.
Walking them lands exactly on the payload end (
28 +blob+24) on every object. - Members after the first begin with 16 bytes of their own index
(
01…01,02…02, …).
| Member | Length | Content |
|---|---|---|
0 (+60) |
560 | The projection record: every fixed offset in the table above |
| 1 | 92 | After 8 bytes that vary between objects (an int32, then 0 or 1), the 16-byte index, 28 zero bytes and eight float64 rDUMMY values: an unused parameter array. Byte-identical from the index onward on 107 of 107 objects |
| 2, 3 | 64 each | The same varying 8 bytes, then the index and 64 zero bytes: identical from the index onward on all 94 (member 2) and 86 (member 3) objects that have them |
| 4, 5 | 72, 258 | Seen once (East_Isa). Member 4 holds ASCII EPSG and int32 28354, the EPSG code of that object's own name, "GDA94 / MGA zone 54". Member 5 begins GDA94. [LIKELY] an authority-code member |
What's still open, deliberately not force-completed:
- The method codes and slot positions of methods other than codes 1, 3,
11 and 14. The reader can still name their values when the registry
holds their text (above), but an object without text stays unnamed.
- The content of members 1-3, which are unset or zero everywhere.
- Not every out-of-range-line_slot blob is projection-related —
scanning ~1700 such blobs in one real file found only 3 with IPJ
content; most instead carry a different tag ("REG ") — see §9 for
what that turned out to be.
- The separate, 4096-byte-stride dictionary region noted elsewhere in
this project (hundreds of generic named projections — US state-plane
zones, etc.) is confirmed to be Geosoft's own bundled reference
catalog, not survey-specific data — a different thing from the
per-database records described here, even though both use IPJ-style
naming and coexist in the same file.
9. The "REG " registry: settings and processing-history log¶
[CONFIRMED] rich real content, on all 3 agencies this project has
files from. The binary framing is partially decoded: a fixed
128-byte preamble common to every REG blob is [CONFIRMED] (1,033
of 1,033 real instances checked); what follows it is [CONFIRMED] to
follow one grammar, entries then nested objects ([CONFIRMED],
below). provenance/notes.md §6.8, §6.8c.
The fixed preamble (§6.4 has the 4670802/REG\0 identity this
starts from): the ordinary 48-byte blob header, its +44 type-code
field holding the object's own name instead; 32 more zero bytes; a
repeating 4-byte constant; a length-like field; then two 16-byte
tag blocks — a separator constant, the FourCC tag "REG ", and two
int32 counts, then the same separator, the tag "VV ", and two more
fields. Every REG blob's first 128 bytes matches this exactly.
provenance/notes.md §6.8c has the full byte-offset table.
The registry grammar -- [CONFIRMED], 1,819 of 1,820 REG objects
(corpus and supplied file). There is one layout, not three forms:
+24 int32 payload length: the object ends at 28 + this
+28 ff 00 f0 0f object frame; its length at +32 ends at +60 + length
+60 ff 00 e1 1e member frame; its length at +64 ends at +92 + length
+92 .. +127 "REG " and "VV " tag blocks (unchanged)
+124 int32 n: the number of key/value entries
+128 n × 256-byte slots, each KEY\0value\0
+128 + 256n int32 m: the number of nested objects
m nested objects, each opening with ff 00 f0 0f and a length
that, like the outer frames, counts from 28 bytes after itself
All three length fields end on the same byte. The measured rules, over 1,820 objects:
- The frame lengths hold on 1,820 of 1,820.
- With
m = 0, the object ends 4 bytes after its last slot. This holds on 1,600 of the 1,601 objects withm = 0. - With
m = 1, the object ends exactly where its nested object ends. This holds on 219 of 219. The nested object was aMAKERrecord every time, naming the GX that made the channel (newchan.gx,linechan.gx,geogxnet.dll(Geosoft.GX.MathExpressionBuilder...)). Its text is UTF-8, begins with a BOM, and ends in a0x1Abyte on the object's last byte. - The single exception is in the supplied file. It has an entry
named
ASSOCIATED.DB_TABLE$$$with an empty value. Its table is stored after the countm = 0as a further"REG "tag block, which this grammar does not describe.
What earlier looked like separate forms:
- "Nested" objects:
n = 0,m = 1, which is why the int32 at+128reads1. - "Empty" objects:
n = 0,m = 0, ending at+132. - "Numeric array" objects and "dirty slot 0" objects: also
n = 0,m = 0. Their bytes lie entirely outside the declared payload, so they are [LIKELY] leftovers of an earlier, longer version of the object. A dirty-slot-0 tell-tale byte at+132is the 5th character of an overwritten key (107 of 110 match a known key). - Key-shaped slots past
n: also outside the payload in the 8 objects that have them.
Support for the leftover reading: dropping these removes every _PJ_*
projection key that had landed on a non-coordinate channel (e.g.
Fiducial, GPS_Height on AG106386). The ones kept sit only on
coordinate pairs. The same "payload length from +28" reading fits
Line Selection objects too (§2.1).
The reader (find_channel_settings) decodes the n entries. It does not
yet expose the nested MAKER records.
Group-class channel lists -- [CONFIRMED]. In files with group lines
(§3.2), __dbreg holds ASSOCIATED.<class> whose value is a
comma-separated list of channel names. Example: ASSOCIATED.DB_TABLE =
dgrf_total,Longitude,Latitude,mag_value_,.... These are the channels
associated with that group class, per the vendor's set_group_class
documentation. A sibling entry ASSOCIATED.<class>$$$ with an empty
value accompanies it. It is an ordinary counted entry, and the registry
grammar above holds for these objects (six USGS files). The one grammar
exception (the supplied file) is additional data after such an entry.
The VV: a vector of fixed-width strings -- [CONFIRMED], 1,985 VVs in
the corpus. Every 00 1a cc ff + VV block is followed by an int32 0,
an int32 element type, an int32 count n, then the n elements. The
element type is negative, and minus it is the element width in bytes:
the same convention as a channel's string type (§3.1).
| Object | Element type | Element |
|---|---|---|
registries (__<n>, __dbreg) |
-256 (the "constant" 00 ff ff ff at +120) |
KEY\0value\0 |
Display List |
-82 (33 objects) or -130 (15, newer files) |
channel name\0handle\0 |
Database Extension Objects (leftover bytes only) |
-256 |
-- |
So a registry's +124 entry count is simply the VV's length. A
Display List ends exactly after its elements.
The administrative-object preamble, generalized -- [CONFIRMED]. A
framed object starts with the object frame at +28 (ff 00 f0 0f,
length), an int32 1 at +40, and the object's class name at +44
(REG, IPJ, EXT, META). Then comes the member frame at +60
(ff 00 e1 1e, length), and at +92 a 00 1a cc ff block carrying the
member's 4-character code (REG, JPI, LMSL, ATEM). A
Display List is a bare VV object instead: a 00 1a cc ff block at
+28, its object frame at +40 and class name VV at +56.
Line Selection has no framing at all (§2.1).
MAKER records: how a channel was made -- [CONFIRMED], 305 of 305.
A registry's nested object (the grammar above) is a MAKER record. In
order, it holds:
- the object frame and class name
MAKER; - the member frame;
- a
00 1a cc ffblock with the codeMAKE, then int321; - int32
L1and the tool that made the channel (L1bytes, NUL included), then a 2-byte field, always 0; - int32
L2and the tool's label; - the tool's parameters as text lines
TOOL.KEY="value", CRLF-separated and ending in0x1A.
The text is UTF-8 with a BOM on 291 records, and plain ASCII without one
in the 2004-2006 files. 24 records have an empty parameter set (just the
BOM and 0x1A). Examples:
newchan.gx("New channel"):NEWCHAN.NAME,DTYPE,ARRAYSIZE,DISPWIDTH,DISPDIG.geogxnet.dll(Geosoft.GX.MathExpressionBuilder...): the formula (CHANNELINPUTBOX="ch_13=ch_5 - ch_12;").lookupdbch.gx: source database and channels.newxy.gx: old and new coordinate channels and projections.
28 distinct tools occur. On all 32 newchan.gx records, NAME,
ARRAYSIZE, DISPWIDTH and DISPDIG equal the channel's own name,
+118, +94 and +96 (§3.1).
Implemented as pygdb.find_channel_makers / GDB.channel_makers
(records attributed by handle, like find_channel_settings) and
pygdb.find_display_lists / GDB.display_lists (each entry's handle
resolved to the channel's current name; the live copy of each object
taken from the blob directory).
__dbmeta: a typed metadata tree -- [CONFIRMED] container, [UNKNOWN]
node grammar. Class META, member ATEM. The member header holds 11
int32s: 2, a count repeated twice (329-363), three constants
(24, 27, 63), two more varying counts, 0, the zlib stream's length,
and the decompressed length. A zlib stream (DB_COMP_SIZE magic, §7.1)
follows. The decompressed 11-12 KB are node records: a level byte
(02/03), a kind letter (F, L, G, D), int32 links (-1 for
none) and a NUL-terminated name. They cover Geosoft's type vocabulary
(Base, Attributes, Types, IPJ Class, META Class, ...) and a
snapshot of the database's own metadata:
- channel names;
LABEL/UNITSvalues;X_Channel Easting/Y_Channel Northing;_PJ_*keys.
The content differs per file. Seen only in three 1991 GSQ files.
A rarer sibling object, tagged "META\0" instead of "REG\0", wraps a
real compressed stream — and it is Geosoft's own internal type library, not
survey data. Same 128-byte preamble, but the first nested tag is "ATEM",
followed by the 16-byte page-primitive magic (§7.1) and a genuine zlib
(DB_COMP_SIZE) stream. Found on 3 real files (all 1991 GSQ TEM/EM
surveys), 4 instances, each decompressing cleanly to 11-12KB of a real,
regularly-framed object graph whose readable strings are Geosoft's own
class/type vocabulary ("IPJ Class", "META Class", primitive types,
attribute names) headed by "Geosoft"/"Core"/"Types"/"Objects" — the
same category of thing as the bundled projection dictionary already
mentioned in §8, not per-survey content. Not decoded field-by-field.
The majority of "reserved/administrative" blobs (§6.4, §8) are not
IPJ records — they start with a different 4-byte FourCC-style tag,
"REG ", using the same general tagged-object convention as IPJ
(48-byte blob header, then a NUL-terminated short name, then further
tagged sub-objects — e.g. a nested "VV " (vector-value) tag,
sometimes itself containing a named sub-field like "CLASS" or
"MAKE"). [CONFIRMED] as a real, reused framing convention, on
USGS, GSQ, and Ontario files alike; [UNKNOWN] for the exact field
boundaries within it.
What it actually contains: Geosoft Desktop's own persistent settings/processing-history registry, recovered by searching for readable strings rather than parsing a byte-exact structure — and independently cross-validated against real sidecar/ground-truth data on every agency checked:
- Real GX tool run records: literal tool identifiers
(
"geogxnet.dll(Geosoft.GX.MathExpressionBuilder.MathExpressionBuilder; RunChannel)"), human-readable names ("Channel Math Expression Builder","Low-pass filter...","New channel","Copy channel"), and persistedTOOLNAME.PARAMNAME="value"parameter strings — including, on a GSQ file, real channel-creation parameters (NEWCHAN.DTYPE="Double",NEWCHAN.DISPDIG="4") that are a useful cross-reference for some of this project's still-unknown channel-record display fields (§3.1). - Real, user-entered processing formulas, referencing this exact
file's own real channels every time. USGS:
ch_9=comp_mag - ch_8; ch_9=ch_9 + 48066.0;(a base-level/diurnal correction — the pairedReadme.txtindependently describes exactly this kind of correction). GSQ: a YYYYMMDD date formula referencing a realDate_channel, and a grid-math formula referencing a real external SRTM elevation grid file. Ontario: formula fragments referencing real channelsmag_diurn,mag_igrf,mag_lev,mag_gsclevel— all verified against that exact file's own channel list. - Real processing dates, on the USGS file, falling inside the survey's documented flight/processing window.
- Provenance labels naming real intermediate working files
(
LABEL="Source: .\delete.gdb",LABEL="Source: .\gps\mag_gps.gdb"). - Per-channel display units, in a compact
UNITS\0<code>,<count>form (dega,1= decimal degrees,m,1= meters) — the first place in this investigation a real channel's display unit has been found at all (the channel symbol table itself, §3.1, has no confirmed units field). - A textual serialization of the same per-database
IPJprojection settings found in binary form in §8, on every agency checked — e.g."NAD83 / UTM zone 11N"(USGS),"GDA2020 / UTM zone 54S"(GSQ),"NAD83 / UTM zone 17N"(Ontario) — plus internal key names (_PJ_NAME,_PJ_ELLIPSOID,_PJ_DATUM_TRANSFORM,_PJ_PROJECTION,_PJ_UNITS,_PJ_X,_PJ_Y,_PJ_IPJ) that plausibly name the binaryIPJrecord's internal sub-fields. On the GSQ file, every one of six numeric parameters in this textual record (6378137,0.0818191910428158,141,0.9996,500000,10000000) matches the paired ASEG-GDF2.prjsidecar exactly, digit for digit — the strongest single confirmation in this investigation. On Ontario, the same pattern gives the correct northern-hemisphere convention (false northing0, not10000000) and the geodetically correct central meridian (-81) for UTM zone 17N. DB_CHAN_X/DB_CHAN_Y/DB_CHAN_Z(matchingDB_CHAN_X=0 DB_CHAN_Y=1 DB_CHAN_Z=2from the vendor's own published source, §2) — a NUL-terminated key immediately followed by a second NUL-terminated string naming the real channel that plays that coordinate role. [CONFIRMED] universal for X/Y across all 22 real files, all 3 agencies (23% for Z); every resolved value checked and found to be a real, exact channel name.provenance/notes.md§6.8b has the full derivation, including a real complication (this format's append-only blob storage can leave stale, differing copies of the same key -- resolved by validating each candidate against the file's own real channel table rather than trusting position). Implemented aspygdb.registry.find_channel_roles/GDB.coordinate_channels, and used byGDB.to_geoh5to pick each line's coordinate channels automatically wherever the registry confirms them.- Real per-channel settings beyond the coordinate roles —
UNITS,LABEL,FORMULA, the_PJ_*projection fields, and whatever else a real file happens to have written — are decoded the same way, by reading the"REG "object's flat key/value framing directly (provenance/notes.md§6.8c) rather than searching for one specific marker. [CONFIRMED] for the framing (245 of 245 clean instances match exactly, §6.8c); decodes a registry's key/value entries, not its nested objects (theMAKERrecords, not yet exposed). Implemented aspygdb.registry.find_channel_settings/GDB.channel_settings. The owning channel comes from the object's blob symbol -- [CONFIRMED]: the object atblob_index = data_slots + kis named by blob-symbol slotk(read bypygdb.read_blob_symbols), and a per-channel object is named__<n>withn − (blobs_max + lines_max)its channel slot (§2.1). On a corpus-wide oracle (aLABELequal to a channel's name) this names the right channel 218 times; the blob index's own remainder (blob_index % chans_max), which the first implementation used, named it 0 times. Objects with a line handle or any other name are skipped.provenance/notes.md§6.2d. Only the first+124slots of an object are read (the entry count, below), so leftover keys from an earlier version of the object are not returned. With that in place,_PJ_*projection keys land only on real coordinate channel pairs across the corpus.
Correction — actually [CONFIRMED] universal across all 22 real files,
not the patterned absence previously documented here. An earlier
draft of this section, based on provenance/notes.md §6.9, claimed
zero REG or IPJ blobs in seven specific files (the three 1991
Questem-era Mount Gordon files, two Melinda Downs magnetic-data files,
and two derived/inversion-output databases). That scan (and the
provenance script it was based on, provenance/scripts/
reg_ipj_full_scan.py) identified "administrative" blobs with a
hardcoded line_slot > 700 cutoff — reasonable for the specific
files it was tuned against, but wrong in general: administrative blobs
in some real files sit at much lower line-slot values (e.g. line_slot
= 200 — suggestively the same value as the vendor's own
DB_CATEGORY_LINE_GROUP=200 constant, §2 — in the Mount Gordon files),
so a fixed 700 cutoff silently skipped them without any warning.
Re-scanning all 22 real files with pygdb.registry.find_coordinate_systems
(and the equivalent raw-tag scan) using a per-file threshold —
every line-table slot beyond that file's own highest real line index,
from pygdb.gdb_reader.read_lines(), rather than one constant — finds
both REG and IPJ content in all 22 of 22 real files, including
every one of the seven previously reported as empty. This is
[CONFIRMED] by direct re-test, not a theoretical fix. The tool-run
records, processing formulas, and projection names described above in
this section are present far more broadly than originally reported;
whether truly every real .gdb file has REG/IPJ content, or some
still don't (e.g. a database that really was never interactively
opened in Oasis montaj), remains open — only that this specific
7-file "some files have none" claim was a scan artifact, not a real
finding.
The two "constants" ff 00 f0 0f and ff 00 e1 1e are frame
markers -- [CONFIRMED] structure, [LIKELY] names. They are the only two
values of the self-checking (0xFF, 0x00, X, ~X) shape anywhere in any
administrative object. Each is followed by an int32 length that counts
from 28 bytes after itself. 0xF0 frames enclose 0xE1 frames:
- one
0xE1frame in a registry; - a chain of 2 to 6 in an IPJ object (§8).
So 0xF0 reads as "object" and 0xE1 as "member". The outer frame
lengths hold on every REG, IPJ, EXT and __dbmeta object in the
corpus.
What's still open: the exact field boundaries inside a member beyond
what the grammar above gives; the $$$-suffixed entry and the
extra "REG " block after it, seen once in the supplied file; and
the MAKER text's own fields. Line-handle objects are per-line
registries with no entries, §2.1); why a
LABEL/UNITS slot specifically is populated or a bare
placeholder — two sibling keys, CLASS and FORMULA, turned out to be
fully deterministic instead (always empty / always populated). Re-run
corpus-wide with the correct channel attribution, LABEL/UNITS
population is not explained by the channel's dtype, array-ness, display
format, or whether the object is the live copy. Data presence cannot
discriminate, because every attributed channel has data. Population is
instead strongly per-file (all or nothing in most files), which points
at the writing tool; provenance/notes.md §6.8c; and whether any real
file has genuinely no REG/IPJ content at all. (The formerly
unidentified third administrative-blob tag on two GSQ files is the
plain-text OE.DB_ACTIVITY_LOG object; every administrative blob is
now identified by its blob-symbol name, §2.1.)
10. Cross-validated across an independent, non-Geosoft format¶
[CONFIRMED], provenance/notes.md §6.6c. A real .gdb/.geoh5 pair for the
same delivery was cross-checked: .geoh5 is Seequent's newer, openly
HDF5-specified successor container, read here via the independent
open-source geoh5py library (Mira Geoscience, LGPL-3.0-or-later —
unrelated to Geosoft's proprietary engine). The set of real survey line
names recovered from the .gdb's own line symbol table (§3.2) matched
exactly — 258 of 258, zero discrepancies either direction — against
the set of named objects independently read from the paired .geoh5
file by a completely unrelated toolchain.
11. Known real-world oddities (recorded, not resolved)¶
These are real, observed, and safe to ignore for correct reading (a
conforming reader already handles or skips all of them defensively),
but their actual meaning is genuinely [UNKNOWN]. Where a single
field value would help (the header words, channel +108/+116, the user
table, line types and the line block, projection method codes, slot 7,
IPJ members 1-3, the MAKER field and the EXT list), the reader issues
a GDBUnseenFeatureWarning when a file differs from every file seen, so
a user can report it (pygdb.unseen):
- The
f0f0f0f0header-signature variant (§2) — seen twice, in two unrelated real deliveries, always the identical 4 bytes. - The user-record path (§3.3): why 11 files have none, and whether the 3 characters hidden under the user name are an ellipsis. (Its layout and truncation rule are now decoded.)
- The line record's 20-byte dummy-filled block (§3.2) — real, confirmed present, but never seen holding anything but the vendor's dummy sentinels. (The fields around it are now explained: they are the previous line's type, flight and version, §2.1.)
- Reserved/administrative blobs with out-of-range line numbers and a
constant
4670802in place of a valid type code (§6.4) — the constant itself is now explained (it is the object's own"REG\0"/"IPJ\0"name, misread as a type code), and §9 now has a partially decoded binary framing for the REG case, but most of a REG blob's content is still recovered by string search, not full decoding. - Parts of the blob-header trailer (§6.3, §7.4):
+24..+31in the plain header.+16(timestamp or unset) and+20(blob class) are now explained. - The channel-record
+108field (§3.1) — always1.0, no confirmed meaning. (+94/+96are now the display width and decimals.) - The superuser record's
0x20000category bit (§3.3). (The line-handle REG objects are now explained as always-empty per-line registries, §2.1.) - Why the directory's
0x4bit alternates between rewrites (§2.2) -- the pattern is decoded, the purpose is not. - The exact reason some whole files/blobs never engage their declared compression mode (§7.6) — size is ruled out; delivery/tool-version provenance is an untested candidate.
- Why a whole channel's blobs were freed (§2.2) — the free list and header word 116 account for every unlisted blob, but not for the action that freed them.
- Many of the other fields of the line, blob-symbol and channel records are
runs of the vendor's dummy values (float32
-1e32= bytesae c5 9d f4, float64-1e32) or small integers with no known meaning; seeprovenance/notes.md§6.1c "What remains unexplained" for the census. - A line-record category code of
65636(§3.2) on a real, named ("L0") but dataless line-table slot — seen on one real GSQ file (rm001141). Now [LIKELY]0x10000 | 100, a freedNORMALline: the same0x10000bit marks exactly the freed slots of the blob-symbol table (§2.1). Still open: that the bit means "freed" is inferred from the blob-symbol table alone, and no second line example exists.
12. Reference implementation¶
A working Python reader implementing everything marked [CONFIRMED] above lives in this repository:
pygdb/gdb_reader.py— header parsing, full symbol-table decode (channels, with VA/array width; lines, §3.2 — note the known indexing caveat there), and the complete blob-index reader:blob_region_start(),iter_blobs(),find_blob(line_slot, channel_slot),read_blob_values()(handles all three compression modes, single- and multi-page, and auto-detects the "bare blob" variant).pygdb/lzrw1.py— the from-scratch canonical LZRW1 decoder (DB_COMP_SPEED), including both the compressed and stored-raw chunk cases.pygdb/grd_reader.py— a fully solved reader for the sibling.grdgrid format (not.gdb, but the same container family, and the first place the shared 16-byte page-primitive magic was found).pygdb/registry.py— best-effort extraction of coordinate-system names from REG/IPJ administrative-blob content (§8-9).pygdb/gdb.py—GDB, a user-facing, name-based wrapper: list lines/channels, see which channels actually have data on a given line (§1's sparse (line, channel) grid), random-access reads by (line name, channel name), and a description of the file's compression mode and coordinate system(s). Corrects the §3.2 line- indexing caveat against the real blob chain before exposing lines by name.rust/src/lib.rs— an optional, opt-in Rust port (pygdb._native, built withmaturin) of this reader's two measured CPU-bound hot paths: LZRW1 decompression and fixed-width string decoding. Purely an accelerator, not a second implementation of the format -- the Python reader above stays the source of truth, and_nativeis validated against it bit-for-bit on this project's real sample corpus, not just synthetic fixtures.
Run python -m pygdb.gdb_reader <path-to.gdb> for a demo: header
fields, the full channel list, and a decoded sample of real data from
the blob chain.
Robustness. All three modules fail gracefully on a blob, chunk, or
record they can't parse — a truncated file (cut-off download, or a
blob chain that runs past EOF), an unrecognized administrative-blob
variant, or anything else that doesn't fit the confirmed structure
above. They return whatever was successfully decoded up to that point
(rather than losing it to an unhandled exception) and issue a clear
GDBParseWarning/GRDParseWarning (Python's standard warnings
module) identifying what couldn't be decoded and why, so a caller can
tell a complete result from a partial one. lzrw1.py raises a single
well-defined LZRW1DecodeError for an undecodable chunk rather than a
bare AssertionError/IndexError. Verified directly against
deliberately truncated/corrupted real files, and against the full
22-file corpus (unchanged output, zero spurious warnings). See
provenance/notes.md section 6.10.
13. What's not covered here¶
Deliberately out of scope for both this document and the underlying
research: writing or mutating .gdb files. Everything above
describes reading an existing file only.
Also not attempted or only partially done: full decoding of the
REG/coordinate-system record format beyond what §8 covers, and full
decoding of the line-table record layout beyond the two fields in
§3.2. See provenance/notes.md's "Natural next steps" section for the current,
prioritized view on what (if anything) is worth pursuing next.