Research Log¶
Historical document
This is the chronological research log from the original
investigation, preserved as-is for provenance. It predates the
project's reorganization into the pygdb package, so paths like
reader/gdb_reader.py and scripts/ refer to that earlier layout
(now pygdb/gdb_reader.py and docs/provenance/scripts/
respectively), and its own cross-references to NOTES.md/LOG.md
mean notes.md and this document.
A running, chronological lab notebook. Each entry records what was consulted,
what was learned (or not learned), and how it feeds into hypotheses about the
.gdb binary format. Polished conclusions live in NOTES.md; this file is
the audit trail showing how we got there. Entries are numbered and dated by
research session, not necessarily by wall-clock time within a session.
Hard constraints in force for this entire log (see task brief / repo README): no Geosoft engine/SDK of any kind ever installed or run.
Session 1 — 2026-09-08¶
1.1 Repo setup¶
Created E:\Repos\pygdb-cleanroom as a brand-new git init repository,
unrelated to any existing history, per the task's hard constraint #4.
1.2 GeosoftInc/gxpy — license check¶
- Source:
https://raw.githubusercontent.com/GeosoftInc/gxpy/master/LICENSE(vendor-published source, read via WebFetch) - Finding: Confirmed BSD 2-Clause License, Copyright 1995-2018 Geosoft Inc.
Standard permissive terms (redistribution with attribution, no warranty).
This clears
gxpyas fair-game source material per the task brief.
1.3 GeosoftInc/gxpy — repo shape¶
- Source:
https://github.com/GeosoftInc/gxpy(repo root, WebFetch) and the GitHub API tree listing (api.github.com/repos/GeosoftInc/gxpy/git/trees/master?recursive=1, fetched via curl, saved locally as scratch data, not committed). - Finding: Two layers:
geosoft/gxpy/*.py— a hand-written, Pythonic convenience layer (gdb.py,vv.py,va.py,utility.py, etc.) that wraps...geosoft/gxapi/*.py— auto-generated thin bindings that call into compiled DLLs shipped in the same directory (geoengine.core.gx_utf8.dll,geogx_utf8.dll, etc.) via agxapi_cy(Cython) layer.- Important implication: none of this Python source can leak
byte-level
.gdbfile-format knowledge beyond docstrings and constants — the actual file I/O happens inside the compiled DLL, which we do not have, do not want, and are not touching. We only read the text of the.pyfiles (fair game, BSD-licensed, published source) — never imported or executed any of it, and the DLLs were never downloaded.
1.4 gxpy source reading — vv.py, utility.py¶
- Source:
geosoft/gxpy/vv.py,geosoft/gxpy/utility.py(raw GitHub content, WebFetch) - Finding:
vv.py(GXvv, the "vector" wrapper used for channel data) confirms the conceptual model — a VV is data +(fid_start, fid_incr)— but again, all real work delegates togxapi.GXVV, i.e. the compiled engine. No byte layout here, but confirms the fiducial (start, increment) addressing model matches what the public docs describe (see 1.7 below). utility.pygave usgx_dummy()andgx_dtype()/dtype_gx()— Python-level dummy-value and type-code lookup tables that referencegxapi.rDUMMY,gxapi.GS_DOUBLE, etc. by name but don't define their numeric values in this file. Sent us looking for the actual constants (see 1.5).
1.5 gxapi/init.py — the constants block (high value)¶
- Source:
https://raw.githubusercontent.com/GeosoftInc/gxpy/master/geosoft/gxapi/__init__.py(downloaded via curl, 7825 lines; vendor-published source) - Finding: This file has a clearly marked
### block Constants # NOTICE: Do not edit anything here, it is generated codesection with literal numeric constants generated straight from Geosoft's C headers. Extracted (all directly quoted, not inferred):These are vendor-published symbolic names and literal values (from reading public source), safe to use directly per the task brief. They are strong hypothesis-generators for byte layout (e.g.iDUMMY = -2147483647 rDUMMY = -1.0E32 GS_S1MX=127 GS_S1MN=-126 GS_S1DM=-127 (signed byte) GS_U1MX=254 GS_U1MN=0 GS_U1DM=255 (unsigned byte) GS_S2MX=32767 GS_S2MN=-32766 GS_S2DM=-32767 (signed short) GS_U2MX=65534 GS_U2MN=0 GS_U2DM=65535 (unsigned short) GS_S4MX=2147483647 GS_S4MN=-2147483646 GS_S4DM=-2147483647 (signed long) GS_U4MX=0xFFFFFFFE GS_U4MN=0 GS_U4DM=0xFFFFFFFF (unsigned long) GS_S8DM=0x8000000000000000 GS_U8DM=0xFFFFFFFFFFFFFFFF (64-bit) GS_R4DM = -1.0E32 (float32 dummy) GS_R8DM = -1.0E+32 (float64 dummy) GS_BYTE=0 GS_USHORT=1 GS_SHORT=2 GS_LONG=3 GS_FLOAT=4 GS_DOUBLE=5 GS_UBYTE=6 GS_ULONG=7 GS_LONG64=8 GS_ULONG64=9 GS_FLOAT3D=10 GS_DOUBLE3D=11 GS_FLOAT2D=12 GS_DOUBLE2D=13 GS_MAXTYPE = 13 GS_TYPE_DEFAULT = -32767 DB_CATEGORY_CHAN_BYTE=0 ... DB_CATEGORY_CHAN_ULONG64=9 (mirrors GS_* order) DB_CATEGORY_LINE_FLIGHT=100 DB_CATEGORY_LINE_GROUP=200 DB_CATEGORY_LINE_NORMAL=100 DB_CHAN_FORMAT_NORMAL=0 EXP=1 TIME=2 DATE=3 GEOGR=4 SIGDIG=5 HEX=6 DB_CHAN_UNPROTECTED=0 DB_CHAN_PROTECTED=1 DB_CHAN_X=0 DB_CHAN_Y=1 DB_CHAN_Z=2 DB_COMP_NONE=0 DB_COMP_SPEED=1 DB_COMP_SIZE=2 DB_GROUP_CLASS_SIZE=256 DB_INFO_BLOBS_MAX=0 LINES_MAX=1 CHANS_MAX=2 USERS_MAX=3 DB_INFO_BLOBS_USED=4 LINES_USED=5 CHANS_USED=6 USERS_USED=7 DB_INFO_PAGE_SIZE=8 DB_INFO_DATA_SIZE=9 DB_INFO_LOST_SIZE=10 DB_INFO_FREE_SIZE=11 DB_INFO_COMP_LEVEL=16 DB_INFO_FILE_SIZE=17 DB_INFO_INDEX_SIZE=18 DB_INFO_BLOB_SIZE=19 DB_INFO_MAX_BLOCK_SIZE=20 DB_INFO_CHANGESLOST=21 DB_LINE_TYPE_NORMAL=0 BASE=1 TIE=2 TEST=3 TREND=4 SPECIAL=5 RANDOM=6 DB_LOCK_NONE=-1 READONLY=0 READWRITE=1 DB_SYMB_BLOB=0 DB_SYMB_LINE=1 DB_SYMB_CHAN=2 DB_SYMB_USER=3 DB_SYMB_NAME_SIZE = 64 DB_ARRAY_BASETYPE_NONE=0 ... ENERGIES=8 (array-channel sub-kinds)DB_SYMB_NAME_SIZE=64suggests fixed 64-byte name fields in a symbol table;DB_GROUP_CLASS_SIZE=256suggests a 256-byte class-name field) but not yet confirmed against real file bytes — flagged as hypotheses in NOTES.md, not facts.
1.6 GXDB.py — page size / compression-level docstrings¶
- Source:
https://raw.githubusercontent.com/GeosoftInc/gxpy/master/geosoft/gxapi/GXDB.py(curl, 5122 lines; vendor-published source) - Finding:
GXDB.create_comp()/create_ex()docstrings (not byte-layout, but authoritative on defaults/semantics):Default:param page: Page Size Must be (64,128,256,512,1024,2048,4096) normally 1024 :param level: DB_COMP (i.e. one of DB_COMP_NONE/SPEED/SIZE, see 1.5)create()(no explicit page/level) semantics: lines=200 max, chans=50 max, blobs=chans+lines+20, users=10, cache=100, super="SUPER", password="". These are creation-time capacity parameters, not necessarily literal on-disk header fields, but a reasonable prior for what a header/superblock records. All actual DB I/O methods (get_info, etc.) bottom out ingxapi_cy.WrapDB._xxx(...)— confirms again there is no byte-layout information obtainable from this file beyond docstrings; it's a pure RPC-style binding to the compiled engine.
1.7 Seequent help.seequent.com — compression & structure docs¶
- Source A:
https://help.seequent.com/Oasismontaj/2023.2/Content/ss/edit_preprocess_data/view_edit_spreadsheet_data/c/database_compression.htm(vendor documentation) - Finding: Three modes: No compression, Compress for speed
(~58% of uncompressed size, ~3x faster r/w), Compress for size
(~19% of uncompressed size, ~25% faster r/w). Benchmarks are
best-case (large DB, mostly doubles). Explicitly states: "the
lossless open source library at zlib.net" is used. This directly
maps to
DB_COMP_NONE/SPEED/SIZE = 0/1/2from 1.5 — vendor doc names the algorithm family (zlib/deflate), a huge constraint on later byte analysis (rules out having to reverse-engineer a bespoke compressor). - Source B:
https://help.seequent.com/Oasismontaj/2023.1/Content/ss/prepare_om/work_with_databases/c/oasis_databases.htm(vendor documentation) - Finding: Conceptual model: elements (byte/ushort/short/long/ float/double/string scalar values keyed by fiducial) grouped into channels (arrays of elements at a fiducial start+increment), channels grouped across lines (one flight/ground line/drillhole each), all lines sharing one global channel-name namespace. States explicitly: "proprietary 3-dimensional-file format architecture... columns stored separately... made up of straight binary data." No byte offsets given, as expected from end-user docs.
- Source C:
https://help.seequent.com/Oasismontaj/2023.2/Content/ss/glossary/database.htm(vendor documentation) — same conceptual model, adds "object-oriented database" framing and reiterates elements/channels/fiducial/lines.
1.8 Geosoft GX Developer wiki (Atlassian, public)¶
- Source:
https://geosoftgxdev.atlassian.net/wiki/spaces/GXD93/pages/103415898/Geosoft+Databases(public SDK wiki, vendor-published but developer-facing) - Finding: Corroborates 1.7 exactly (lines/channels, format dates to 1992, "de facto standard" framing, all lines share channel defs). Nothing beyond what the end-user help already gave us — useful as independent corroboration of the conceptual model, not new byte-level data.
1.9 Loop3D/geosoft_grid — sibling-format reader (high value)¶
- Source:
https://raw.githubusercontent.com/Loop3D/geosoft_grid/main/README.mdandhttps://raw.githubusercontent.com/Loop3D/geosoft_grid/main/grd2geotiff.py(curl; MIT-licensed, independent third-party clean-room-ish reader for the sibling.grdgrid format, not.gdbitself) - Finding (README): Explicit credited correction: "compression is
in fact zlib not LZRW1 regardless of what COMP_TYPE says" (credited to
Evren Pakyuz-Charrier). This is independent, non-Geosoft, real-file-tested
confirmation that the on-disk
COMP_TYPE-style enum field can claim a legacy codec name (LZRW1) while the actual bytes are zlib/deflate. Directly corroborates the Seequent doc's "uses zlib" statement (1.7A) for the grid format, and is a strong prior — not yet proof — that.gdb's own compressed pages will likewise be zlib regardless of what any internal enum says. - Finding (grd2geotiff.py full source, read in full):
.grdhas a fixed 512-byte header, parsed witharray.arrayat fixed byte offsets: bytes 0-19 = 5 int32 (ESelementsize/flags,SFsign_flag,NE,NV,KXordering), bytes 20-59 = 5 float64 (spacing/origin/rotation), bytes 60-75 = 2 float64 (z-scaling), bytes 140-183 = assorted int32/float32 optional params.- Per-type-size dummy values are literal small constants (
-127,255,-32767,65535,-2147483647,4294967295,-1e32) — these exactly match theGS_*DMconstants pulled from gxapi/init.py in 1.5, which is a genuine independent cross-check: two unrelated public sources (Geosoft's own generated constants, and a third party's from-scratch reverse-engineering of a sibling format) agree on the dummy-value scheme. Strengthens confidence these dummy conventions are shared across the whole Geosoft container family, including.gdb. - Compressed-grid block layout (their own comment: "There is an
unexplained 16 byte header that we also need to remove"): after the
512-byte header, a compressed grid stores
n_blocks(int32 @ offset 8),vectors_per_block(int32 @ offset 12), then an array ofn_blocksint64 block-start-offsets, then an array ofn_blocksint32 compressed-block-sizes, then the zlib-compressed blocks themselves (each block independentlyzlib.decompress-able). This offset-table + size-table + independently-compressed-chunks pattern is a concrete, real, working example of how a Geosoft container implements paged/blocked zlib compression. Not proof.gdbuses the identical layout, but a strong structural hypothesis given.gdb's ownDB_INFO_PAGE_SIZE/DB_INFO_MAX_BLOCK_SIZEconstants (1.5) imply a similar page/block model.
1.10 Sample-file hunt — status: not yet successful, several dead ends logged¶
Extensive search for downloadable real .gdb files from government open-data
portals. Recording dead ends explicitly per the "log what didn't work too"
instruction:
- Ontario GeologyOntario (
geologyontario.mndm.gov.on.ca): Found metadata pages for GDS1089 / GDS1251 (both explicitly say they include Geosoft.gdbformat). The old direct-download endpoints (.../mndmfiles/pub/data/records/GDS1089.html,.../mndmaccess/mndm_dir.asp?type=pub&id=GDS1089) now 404 or redirect to a JS-driven "hub" site (hub.geologyontario.mines.gov.on.ca) that a plain WebFetch/curl could not extract a file listing from in the time spent. Not abandoned, just deferred — candidate to revisit if time allows. - NRCan GDR (
gdr.agg.nrcan.gc.ca): This domain, which manyopen.canada.caCAGDB dataset pages link to as their actual download endpoint (e.g....index-eng.php?data_file_name=Southeast+Manitoba.gdb), is entirely unreachable from this environment — connection refused / timeout on both http and https, confirmed with both WebFetch and raw curl. Given this affects essentially all Canadian CAGDB compilation downloads, this whole avenue is closed off for this session. - NWT Geological Survey (
nwtgeoscience.ca): Found NWT Open File 2015-02 ("Historical aeromagnetic surveys over the north shore of the East Arm, Great Slave Lake, NWT") via its readme PDF (https://www.nwtgeoscience.ca/sites/ntgs/files/nwt_open_file_2015-02_readme.pdf, read in full) — excellent candidate in principle: 11 small 1990s-era surveys (Storimin, Wiscan, NP_Tete, NP_MacKay, NP_Denis, NP_Barnston, MG_Misty, HF_West, HF_Northwest, HF_East, HF_Central), and the readme states explicitly "For each of the surveys, one geosoft-format grid file (GRD) and one geosoft-format database (GDB) were donated" — a perfect paired GDB+GRD set, and old/small enough to likely be a genuinely small download (unlike modern high-density surveys, see below). However, the actual file download is gated behindapp.nwtgeoscience.ca, an ASP.NET WebForms app with__VIEWSTATEpostback state that couldn't be scripted with plain curl in the time available. Deferred, not abandoned. - USGS ScienceBase (
sciencebase.gov): This portal works well — its plain JSON API (sciencebase.gov/catalog/item/<id>?format=json&fields=files) directly lists file names/sizes/URLs, no scraping needed. Surveyed ~15 USGS EarthMRI airborne magnetic/radiometric survey items this way. Every one of them is a modern (2016-2025) high-line-density survey whose combined Geosoft-database zip is huge: smallest found so far isgpr2025_003_databases_geosoft.zipat 817 MB (Kaiyuh Mountains, Alaska, item698bcd98b66b01469f9c9eeb), most are 1-7 GB. Per the task's "avoid multi-GB, prefer smaller" guidance, none of these are ideal for a full download, though their much smaller*_grids_geosoft.zipcompanions (24-185 MB, containing.grdnot.gdb) would be fine sibling-format cross-checks. - Geoscience BC (
geosciencebc.com/projects/2017-sea02/): Direct links found and reachable (GBCR2018-02-GDB-Magnetics.zip~3.3 GB,GBCR2018-02-GDB-Radiometrics.zip~889 MB) but both too large. - New Brunswick (
gnb.ca): Geophysical open data is GIS-service-only (WMS/WCS/ArcGIS REST,.ersrasters) — no.gdbdownloads found.
Next step: Either (a) crack the NWT ASP.NET app or the Ontario hub site for a genuinely small legacy file, or (b) use HTTP Range requests against one of the large-but-reachable USGS ScienceBase zips to pull just the header/index bytes remotely without a full multi-GB download, reserving a full small download for whichever small source pans out.
Session 1 (continued) — sample-file breakthrough + byte-level analysis¶
Note on process: mid-session, git commits started failing (this machine's
global ~/.gitconfig requires GPG-signed commits and no secret key is
available in this environment). Flagged to the operator; decision was to
not bypass signing or touch git config, and just keep accumulating the
trail in this log file (committing to be sorted out later). So from here
on, "logged" means "written to this file," not "committed" — the numbered
entries below are still in strict chronological order of the actual work.
1.11 Loop3D/geosoft_grid own test fixtures (per operator suggestion)¶
- Source: GitHub API tree listing for
Loop3D/geosoft_grid(curl toapi.github.com/repos/Loop3D/geosoft_grid/git/trees/main?recursive=1) - Finding: The repo ships
test_data/test_compressed.grd(262,598 bytes) andtest_data/test_uncompressed.grd(287,656 bytes) — the same grid, saved both ways — plus matching.gi(32,768 bytes each, constant size, probably a fixed-size legend/histogram sidecar) and.xml(9,334 bytes each, GX metadata/projection sidecar) files for both. MIT-licensed, small, downloaded in full tosamples/loop3d_grd_test/. This pair is a self-contained, oracle-free ground truth: same data, two compression states, from a real third-party project — ideal for diffing.
1.12 Deliberate scope decision: excluding GeosoftInc/gxpy's own binary CI fixtures¶
- While re-examining the gxpy repo tree (1.3), noticed it also contains
binary test fixtures/CI outputs under
geosoft/gxpy/tests/(e.g.little.zip,dem.zip,dem_small.zip, and atests/results/tree of.map/.xmlfiles that are clearly artifacts of Geosoft's own CI actually running their engine to produce reference outputs for their test suite). - Decision: did not download or use any of these as ground truth. Even though merely downloading a published file isn't literally "running the engine," using engine-generated binary fixtures as a reverse-engineering oracle defeats the purpose of the clean-room exercise as clearly as running the engine ourselves would — the task's constraint is about not leaning on Geosoft's proprietary engine as an oracle "in any form." Reading their source code (text) is fair game (explicitly source (a)); using their engine's binary output as a reference is not, even if it's sitting in the same public BSD-licensed repo. Recorded here so the boundary is explicit and auditable, not just assumed.
1.13 Zenodo/figshare/PANGAEA — no usable results¶
- Zenodo's public search API (
zenodo.org/api/records?q=...) timed out repeatedly (HTTP 504 / connection failures) from this environment — couldn't get a result either way. Not pursued further given time budget. - Web search for figshare/PANGAEA turned up general commentary that
Geosoft's binary format is not typically what gets archived in these
academic repositories (published supplementary data tends to be
ASCII/CSV/NetCDF, precisely because
.gdbisn't an open standard) — no direct.gdbsample hits. Logged as a dead end, not a contradiction of anything.
1.14 Geological Survey of Queensland (GSQ) Open Data Portal — CKAN API works, downloads blocked by WAF¶
- Source:
https://geoscience.data.qld.gov.au/data/api/3/action/package_search?q=...(public CKAN Action API, no auth) — per operator tip that this is a standard CKAN deployment. - Finding: The search API works perfectly and is a clean, generic way
to query — e.g.
q=geosoft+gravityreturned 4 real exploration-company report packages, several with resources literally named "IP GEOSOFT DATA" (cr035932,cr035272— old (2003) small mineral exploration reports, exactly the kind of small legacy file we wanted). - Dead end: the actual file bytes are hosted on
gsq-prod-ckan-horizon-public.s3.ap-southeast-2.amazonaws.com. Direct S3 URLs returnAccessDenied(need a signed URL). The portal's own/data/dataset/<pkg>/resource/<res>/download/<file>proxy path — which should mint that signed redirect — instead returns an AWS WAFx-amzn-waf-action: challenge(HTTP 202, empty body): a bot-detection JS challenge that plaincurlcannot solve. No workaround attempted (would require a real browser/JS engine, out of scope for this exercise). Logged as a dead end specific to downloading from GSQ; the search API itself remains a good technique, credited to the operator's tip.
1.15 USGS ScienceBase — the working sample-file pipeline¶
- Source:
https://www.sciencebase.gov/catalog/items?q=<text>&format=json(search) andhttps://www.sciencebase.gov/catalog/item/<id>?format=json&fields=files(per-item file listing) — plain public JSON APIs, no auth, no scraping. - Searched
"airborne magnetic radiometric survey"(89 hits) and inspected file listings for ~20 USGS EarthMRI-program items. Every modern (2016-2025) survey's combined Geosoft-database file is large (hundreds of MB to several GB) because it holds the entire multi-thousand-line-km survey at high sample rate — inherent to the data, not a portal limitation. Logged the sizes for ~15 items in case a smaller one is wanted later (see earlier table in the conversation / can be regenerated from the API easily). - The one that worked well: item
5e7be5eee4b01d50927301b3, "Airborne magnetic and radiometric survey of the southeast Mojave Desert, California and Nevada" (Ponce, D.A., and Drenth, B.J., 2020, U.S. Geological Survey data release, DOI: 10.5066/P9UWYYK9). Files: Magnetic_Data.gdb— 777,859,072 bytes — real Geosoft databaseMagnetic_Data.csv— 893,523,364 bytes — the same data as plain ASCII, perReadme.txt: "Aeromagnetic database in Geosoft database format. Contents are described in the survey report" (csv) vs "Aeromagnetic database in Geosoft format" (gdb) — an explicit paired GDB+ASCII export, exactly the source-(f) ground-truth pairing the task brief calls out.Radiometric_Data.gdb— 706,635,776 bytes,Radiometric_Data.csv— 136,762,283 bytes — same pairing for the radiometric channel set.TMI.gxf,K.gxf,eTh.gxf,eU.gxf— ASCII Grid eXchange Format grids (public, documented ASCII format, not Geosoft.grd, but a bonus sibling product) — not downloaded (out of scope, csv+gdb pair was the priority).Metadata.xml,Readme.txt— read in full, gave the exact citation, survey parameters (17,277 line-km, 200 m line spacing, Dec 2019-Mar 2020, flown by EDCON-PRJ Inc.), and a description of every file.- Given the operator's explicit relaxation of the "prefer small" guidance
for well-provenanced pairs, downloaded both
.gdbfiles in full (via plaincurl, foreground-in-background-tool-call — first attempt using a shell&background job got orphaned/killed early by the harness, second attempt running curl directly as the backgrounded tool call worked correctly) tosamples/usgs_mojave_2020/. Both downloads verified byte-exact against the sizes reported by the ScienceBase API. - The two CSVs were not downloaded in full (900MB/137MB would push
total download past a sane amount for what's needed) — instead pulled
a prefix of each via
curl ... | head -c 5000000 > ...(relying on the pipe closing to kill the transfer onceheadhas enough — this works because the endpoint does not honorRange:requests, confirmed below, so a true partial GET wasn't available). This gave real column headers and several thousand real rows of ground truth for both databases:- Magnetic CSV columns:
lat,lon,x,y,gps_elev,radar,dem,drape,fid, raw_mag,comp_mag,base,line,date,flight_number,heading,gps_elev88, calc_radar,time,diurnally_cor_leveled_mag,filtered_base, diurnaly_cor_mag,igrf_correction,final_mag(24 columns) - Radiometric CSV columns:
epoch,calc_radar,dem,fid,x,y,uth,thk,uk, flight_number,date,line,time,Dose,cosmic_raw,upward_counts, gps_elev88,gps_elev,lat,lon,radon,pressure,temperature,humidity, heading,k_raw,STP,tc_raw,th_raw,u_raw,ucorr,kcorr,tccorr,thcorr(33 columns) - First CSV row (magnetic):
lat=34.94592984, lon=-115.28992715, x=656157.595777467, y=3868381.99450346, ..., fid=577342, raw_mag=47657.635, comp_mag=47662.5705, ..., line=L1000, date=2020/01/15, ...— these exact values become the ground truth used below to locate real data inside the real.gdbfile.
- Magnetic CSV columns:
- Range-request note: tested
curl -H "Range: bytes=0-1023"against the ScienceBase file-get endpoint; server ignored the header and returned a full200 OKwith the entire file rather than a206 Partial Content— confirmed this endpoint does not support partial GETs, so the "range-request on a large remote file" fallback plan from earlier isn't available here; full downloads (or the pipe-truncation trick for read-only prefix sampling) were the only options.
1.16 .gdb header — byte-level hypothesis testing against the two real files¶
All of this is my own byte analysis of the two files from 1.15
(Magnetic_Data.gdb, Radiometric_Data.gdb), cross-checked against each
other and against the vendor-published constants from 1.5/1.6.
- Hypothesis (from nothing but curiosity): first bytes are a magic
number. Test:
od -A d -t x1z -v -N 512 Magnetic_Data.gdb. Result — CONFIRMED, both files: first 16 bytes identical in both:21 43 42 44 00 00 00 00 00 00 02 10 08 01 00 00. Bytes 0-3 read as ASCII"!CBD". Bytes 4-15 are a second fixed sub-block, identical across both files — evidently a stable format/version signature, not data-dependent. Status: confirmed magic (16 bytes), meaning of bytes 4-15 beyond "constant" unknown. - Hypothesis: the vendor-documented default DB creation parameters
(
GXDB.create()docstring, 1.6: page size, chans/lines/blobs/users max, page size 64-4096 normally 1024) appear literally as header fields. Test: scanned int32 words in the first 128 bytes of both files, looking for the value 1024 (page size) and other round numbers. Result — PARTIALLY CONFIRMED: - Offset 100 (int32) = 1024 in both files → high-confidence page size field (matches "normally 1024" default exactly, and is a suspiciously specific constant to appear by coincidence).
- Offset 24 (int32) = 50 (Magnetic) / 100 (Radiometric). Hypothesized chans_max (default is 50 per the docstring; Radiometric evidently created with a raised capacity). This one gets a strong independent confirmation in the next entry, not just a guess — promoted to high confidence.
- Offset 40 (int32) = 10 in both files → matches the default
users=10parameter exactly. Hypothesized users_max. Not independently cross-checked beyond matching the doc default in both files (weaker than the chans_max confirmation, but consistent). - Offsets 28, 32, 36, 44, 48, 52, 56, 60, 64, 68, 72, 76, 80, 84, 88, 92, 104, 108, 112 all hold plausible-looking round-ish integers (1070/1200, 100/5000, 1000/1000, 51180/105698, 50000/100000, ...) that differ in patterned ways between the two files but whose exact semantics were not pinned down — flagged as unconfirmed/open in NOTES.md. Offset 104 (580000 / a similarly large value) is a plausible index/table-size field (roughly consistent with where the symbol table sits, see below) but not proven precisely.
- Hypothesis: there's a symbol/name table somewhere holding channel
names, since the format is documented (1.7) as exposing named channels.
Test: searched each file's raw bytes for the literal ASCII channel
names pulled from the real CSV headers in 1.15 (
lat,lon,raw_mag, etc. for Magnetic;epoch,uth,kcorr, etc. for Radiometric). Result — CONFIRMED, and this is the single best result of the session: - All channel names are found clustered together in one region of the file (~572,300-580,700 in Magnetic_Data.gdb).
- The names recur at an exact, constant 128-byte stride (confirmed
by finding
latat 572312,lonat 572440,xat 572568,yat 572696 — each exactly +128 from the last — then verified across all 24 known channels plus, once the full table was scanned mechanically in 128-byte steps from the first hit, 9 more channels not present in the CSV export: a secondtimeentry,__X,__Y,year_jd, and evidently-abandoned working channels literally namedcrap,crap2,deg,ch_11— real leftover mess from whatever processing produced this file, which is itself a nice authenticity signal that this is genuine field data, not a synthetic fixture). - Table layout per 128-byte record (Magnetic file, byte offsets
relative to record start):
+0..+7: 8 bytes, zero in every record seen (reserved / unused in these files — could be a pointer field that's simply unpopulated for simple flat channels; not tested against an array channel since none exist in this dataset)+8..+71(approx): NUL-terminated/NUL-padded channel name (64 bytes budgeted; matchesDB_SYMB_NAME_SIZE = 64read directly fromgxapi/__init__.pyin 1.5 — nice independent confirmation that a vendor-published constant predicted a real structural detail before we found it)+84(int16, not int32 — confirmed by the string-channel test below): data type code. Positive values matchGS_*constants from 1.5 exactly (5=GS_DOUBLEfor every plain numeric channel seen —lat,x,raw_mag,flight_number, etc. are all stored as float64 regardless of the values' apparent integer-ness, e.g.flight_numberholding small integers like 18 is still typed double). Negative values appear for the two text-like channels:line→-64,date→-10. Tested against real string widths: the realdatevalues in the CSV are formatted"2020/01/15"— exactly 10 characters — matching-10exactly (not-10*4as the Python-sidegx_dtype()inutility.pywould compute for a UTF-8-safe VV allocation — confirms the raw on-disk convention is the simpler "negative value = string byte width," and the*4inflation in 1.4 is a Python/VV-layer allocation detail, not an on-disk one).line→-64is consistent with a fixed 64-byte reserved name-sized field for line labels (values seen are short, e.g."L1000", well under 64). Status: high confidence, confirmed by an exact string-length match on real data, a genuine falsifiable test that passed.+92(int16): format code.date's record has3here, matchingDB_CHAN_FORMAT_DATE = 3from 1.5 exactly.time's record has2, matchingDB_CHAN_FORMAT_TIME = 2exactly. Every plain numeric channel has0(DB_CHAN_FORMAT_NORMAL). Status: high confidence, two independent exact matches against vendor-published constants.+94(int16): varies per channel (10-24 range seen) in a pattern suggestive of a decimal-places/significant-digits display setting (e.g.lat/lon= 12, plain elevation-like channels = 10, heavily-corrected derived channels likediurnally_cor_leveled_mag= 24). Status: plausible guess, unconfirmed — didn't find an independent way to verify the exact semantics; noted as open in NOTES.md.+96(int32): small integer (0-5 range seen), no confirmed meaning found. Status: unconfirmed/open.+112(8 bytes): saw00 00 f0 3ftextually in early hex dumps but a cleanstruct.unpack('<d', ...)at this offset does not read as a tidy value once the following bytes are included properly — not resolved, flagged open rather than guessed at.
- Independent cross-file confirmation of the whole table model:
searched for the literal string
SUPER(the default super-user name from theGXDB.create()docstring defaultsuper="SUPER", 1.6) in both files. In both files,SUPER's record sits exactlychans_maxrecords after the channel table's first record — i.e.record_of_SUPER − chans_max × 128 == channel_table_start, and the channel name sitting at that computed table-start offset is the real first column of that file's CSV (latfor Magnetic,epochfor Radiometric — verified programmatically, not by eye). This independently confirms three things at once, in two unrelated real files: (a) offset-24 really is a channel-table capacity (chans_max), (b) the channel table occupies exactlychans_maxfixed 128-byte slots, immediately followed by a differently-typed table (users) starting with the literal default superuser name — matchingDB_SYMB_CHAN=2immediately precedingDB_SYMB_USER=3in the vendor's own enum ordering from 1.5. This is the strongest result of the session: doc says a default name ("SUPER") → hypothesis about table adjacency → falsifiable arithmetic test → passed identically on two independent real files. - Hypothesis: there's a similar table for lines (survey lines), since
linevalues like"L1000"appear as data. Test: searched Magnetic_Data.gdb for the literal bytesL1000(the real first line name from the CSV). Result — PARTIALLY CONFIRMED: found at absolute offset 444,320, and scanning nearby for repeats of a regex[A-Z][0-9]{3,5}\x00(line-name-shaped tokens) found 631 matches, each exactly 128 bytes apart — same table stride as the channel table, reused for a different symbol type, consistent withDB_SYMB_LINE = 1being a sibling ofDB_SYMB_CHAN = 2in the same unified symbol-table scheme. However, the internal record layout differs: the line name sits at relative+32(not+8as in channel records), and a field at relative+108reads100— matchingDB_CATEGORY_LINE_NORMAL = 100from 1.5 exactly (another vendor-constant hit). Not resolved: reconciling the line table's start offset with the channel table's start offset via thelines_maxguess from offset-28 didn't cleanly round-trip (off by a non-multiple-of-128 amount), so the exact arithmetic connecting "header capacity fields" → "table extents" isn't fully nailed down for the line table the way it was for the channel table. Flagged open. - Hypothesis: actual channel data (the float64 series of
measurements) might be findable directly, uncompressed, if this
particular database happened to be created with
DB_COMP_NONE. Test: searched Magnetic_Data.gdb for the raw little-endian float64 bytes of four real values from the first CSV row:fid=577342(as577342.0),raw_mag=47657.635,comp_mag=47662.5705,base=48076.9751. Result — CONFIRMED, and structurally informative: all four values found as literal byte sequences, no decompression needed, at offsets 698416 (fid), 721968 (raw_mag), 745520 (comp_mag), 769072 (base) — each exactly 23,552 bytes after the previous one. This is strong evidence for column-major storage: each channel's values for a run of data (a line, or part of one) are stored as one contiguous run of same-typed values, not interleaved row-by-row — matching the Seequent doc's own language from 1.7A ("columns stored separately") and the general VV/vector model from 1.4/1.7.23552 / 8 = 2944values, plausibly the sample count of the first line/segment. Not yet located: the index/pointer structure that maps a given (line, channel) pair to its byte offset and length in the file — didn't find it in the time available this session. This is the main open item for anyone continuing this work: without it, a reader can validate the format's shape but can't yet do general random-access reads of arbitrary line/channel data.
1.17 Reader implementation + a bonus confirmation from testing it¶
Wrote reader/grd_reader.py (the fully-solved .grd format from §1.9/4)
and reader/gdb_reader.py (header magic + symbol-table walk for .gdb,
per §1.16). Ran both against the real sample files as a final test, not
just during derivation:
grd_reader.pyagainst bothtest_compressed.grdandtest_uncompressed.grd: decoded shape 286×251, identical header fields, and identical decoded values between the two (compressed decompresses to exactly the same floats as the uncompressed original). Confirms the implementation, not just the manual byte-math done earlier, is correct end to end.gdb_reader.pyagainstMagnetic_Data.gdb: found all 33 channels (24 real + 9 leftover/hidden ones) with correct names, matching the manual analysis in §1.16 exactly.gdb_reader.pyagainstRadiometric_Data.gdb: found 58 channels (not 33 — this file'schans_maxis 100 vs Magnetic's 50, and it has more real + leftover channels). This produced an unplanned but valuable additional confirmation: the Magnetic file happens to type every numeric channel asGS_DOUBLE, which left it ambiguous whether the+84type-code decoding was really general or just happened to always see the same value. The Radiometric file has a much richer type mix —ISPD/ISPUdecode asGS_USHORT,flight_numberdecodes asGS_LONG(notably different from Magnetic'sflight_number, which isGS_DOUBLE— same channel name, different real on-disk type in a different file, which is exactly the kind of case that would break a reader that hardcoded assumptions), and a large group of radiometric calibration channels (radon,pressure,temperature,humidity,k_raw,STP,tc_raw,th_raw,u_raw,ULEVL,UPULEVL,URADREF,KRADREF,THRADREF,TCRADREF,UPURADREF,UHOLD,THOLD,KHOLD,UTHRATIO,UKRATIO,THKRATIO) decode asGS_FLOAT. Every decoded value fell inside the knownGS_*enum from §1.5 with no out-of-range/garbage cases — raised the type-code field from "confirmed on one dominant value" to "confirmed across at least four distinct real type codes (GS_DOUBLE,GS_LONG,GS_USHORT,GS_FLOAT) in two independent files." Also noticed this file has two channels literally namedradon(indices 11 and 39, different types) — more real-world leftover-channel mess, consistent with genuine field data.
1.18 Session wrap targets¶
Given the above, the plan for the rest of this session is: write up
NOTES.md distinguishing clearly-confirmed / cross-validated findings
from single-file observations from open unknowns; build a small Python
reader implementing what's solid (magic check, symbol-table walk →
channel names + dtypes, which is genuinely useful on its own and 100%
verified against two independent real files) with the data-block indexing
left as an explicit TODO/open question rather than a guess dressed up as
code; and write the final summary section.
Session 2 — pressure test + VA/array-channel question (2026-09-09)¶
Prompted by two new asks: (1) chase the "array-per-cell" (VA channel) question left open in Session 1, using new real GSQ (Geological Survey of Queensland) sample files the operator downloaded manually through a browser (the GSQ CKAN portal's automated download is behind an AWS WAF bot-challenge, identified in Session 1 §1.14 -- a human clicking through a browser is not a different kind of source, just a different mechanism for fetching the same public data); (2) pressure-test Session 1's [CONFIRMED] claims against files from a completely independent source (different agency, decades, countries, survey vendors) to see if anything breaks.
Same ground rules as Session 1 throughout: no Geosoft software ever run; git commits still blocked (GPG signing, no key -- not bypassed, just kept accumulating in this file per the standing decision).
2.1 Inventory and extraction¶
- Source:
E:\Repos\pygdb-cleanroom\samples\GSQ_Data\*.zip(operator-provided, originally from GSQ's Open Data Portal,geoscience.data.qld.gov.au) - Extracted the four priority archives (skipped the two multi-GB ones,
Kamilaroi.zipandGeorgetown-AGSO.zip, per the operator's "lower priority/skip" note): Holroy-River.zip→em000293/(GEOTEM system, 1992):DB_EM_293.gdb(12,370,944 bytes),DB_Mag_293.gdb(2,638,848 bytes)Scrubby-Knob.zip→em000833/(GEOTEM, 1993):DB_EM_833.gdb(12,444,672 bytes),DB_Mag_833.gdb(4,077,568 bytes), plus real gridded per-gate productsEM_Ch2_833.grd,EM_Ch5_833.grd,EM_Ch10_833.grdMount-Gordon.zip→em001003/(Questem system, 1991 -- a different vendor/system than GEOTEM):DB_EM_MountGordon_1003.gdb(10,800,128 bytes),DB_Mag_Elaine_1003.gdb(2,249,728 bytes),DB_Mag_MountGordon_1003.gdb(50,249,728 bytes), plusEM_Ch4_1003.grd,EM_Ch11_1003.grdNorthern Georgetown_Conductivity.zip→.dat/.dfn/.des/.prj(extracted everything except the ~1GB.dat, which was instead stream-read in place from inside the zip without full extraction -- see §2.7)- Also present, not re-extracted but analyzed in place read-only:
AG106386_Northern Georgetown_Conductivity.gdb(317,259,776 bytes, already sitting extracted insamples/GSQ_Data/-- the operator's lead for the "third compression scheme", see §2.6-2.9). Never written to at any point this session -- all analysis was read-onlyopen(path, 'rb'), so no backup copy was needed (nothing to lose).
2.2 Pressure test round 1: magic + header + channel table on 7 new files¶
- Test: ran the Session-1 header-field extraction (magic bytes,
chans_max/users_max/page_sizeat offsets 24/40/100) against all 7 new.gdbfiles. - Result: the 4-byte
"!CBD"magic [CONFIRMED] on 7/7 (9/9 counting the two 2020 USGS files from Session 1) -- now verified across three different TEM system vendors (GEOTEM, Questem) and decades (1991-2020) from two unrelated agencies (USGS, GSQ). - One real deviation found:
DB_Mag_Elaine_1003.gdbhas bytes 8-11 =f0 f0 f0 f0instead of the usual00 00 00 00seen in every other file (8/9). The rest of that file parses completely normally (sanechans_max/users_max/page_size, a clean channel table). No explanation found for thef0f0f0f0value -- logged as [UNKNOWN], not force-fitted to a theory. Reader updated to treat only the 4-byte magic as a hard requirement and flag (not reject) a differing extended signature. chans_max/users_max/page_sizefields all decoded to sane, plausible values (chans_max20/20/30/20/50/50/50 across the seven files, scaling sensibly with each survey's real channel count;users_max=10 andpage_size=1024 in all seven, matching the documented defaults again).
2.3 Pressure test round 2: the SUPER-anchor structural proof needed a real revision¶
- Test: ran Session 1's
find_channel_table()(locate the literal string"SUPER", walk backchans_max*128bytes) against the same 7 files. - Result: 3 of 7 failed outright (
DB_Mag_833.gdb,DB_EM_MountGordon_1003.gdb,DB_Mag_MountGordon_1003.gdb--ValueError: could not find 'SUPER'), and a 4th (DB_Mag_Elaine_1003.gdb) failed for an unrelated reason (bad initial 4MB read-window size on this larger file, fixed alongside). This is exactly the kind of "needs revision" result the operator asked to be recorded honestly rather than glossed over. - Root-cause investigation on
DB_Mag_833.gdb(chosen because the literal bytesSUPERdon't appear anywhere in the file --data.find(b'SUPER')returns nothing at all, ruling out a mere windowing bug): - Built a generic, anchor-free scanner: bucket every offset
pwheredata[p+8:p+72]decodes as a clean NUL-terminated printable ASCII name byp % 128(the record stride established in Session 1), then look for phases with long runs of consecutive hits exactly 128 bytes apart. This is a strictly weaker assumption than "there's aSUPERrecord" -- it only assumes the 128-byte channel-record stride itself, which was the actually load-bearing Session-1 finding. - This surfaced the real channel table at absolute offset 972012 with
channels
Easting_AGD66, Northing_AGD66, FID, MAG, RADAR_ALT(5 real channels amongchans_max=20slots) -- a sane, expected-looking magnetics database. - Also surfaced, along the way, a large embedded blob region
(roughly offset 316000-343000) containing what is very obviously
Geosoft's own bundled list of named map-projection/coordinate
systems ("SPCS83 Alabama East zone (meters)", "SPCS83 California
zone 5 (US Survey feet)", etc. -- dozens of standard US state-plane
zone names) plus registry-style entries with a double-underscore
prefix (
__dbreg,__2620,__5120,IPJ_X:Y,IPJ_Easting_AGD66:Northing_AGD66) at a different stride (4096 bytes, not 128). This is clearly aDB_SYMB_BLOB-type region (recallDB_SYMB_BLOB=0precedesDB_SYMB_LINE=1/DB_SYMB_CHAN=2in the vendor's own enum, Session 1 §1.5) holding shared reference data (projection dictionary, IPJ objects) rather than per-database content -- interesting, but a rabbit hole relative to the actual goal, so not chased further. Recorded here so the "GEOTEMCHANNEL1" etc. strings that turn up inside this blob (a processing-history record apparently listing every channel name used anywhere in the project, including from the file's own EM sibling database) aren't mistaken for a second copy of the real channel table by anyone re-deriving this later -- they very nearly were mistaken for exactly that during this investigation. - Root cause identified: manually inspecting the real user-table
record (computed as
channel_table_start + chans_max*128= 972012+2560 = 974572) shows the literal bytes...su per\x00 8\x00 3\x00 3\x00 \\x00 s\x00 c\x00 r\x00 u\x00 b\x00 b\x00 y\x00 k\x00 n\x00 o\x00 b\x00 \\x00 D\x00 B\x00 _\x00 M\x00 a\x00 g\x00 _\x00 8\x00 3\x00 3\x00 .\x00 g\x00 d\x00-- i.e. lowercase"super"(plain ASCII, NUL-terminated as expected), immediately followed by what decodes as UTF-16LE text reading"833\scrubbyknob\DB_Mag_833.gd[b]"-- an embedded Windows file path, apparently the database's own creation path, stored right after the short ASCII username field within the same record. (This second, UTF-16 sub-field was not previously characterized in Session 1 and is logged here as a genuinely new, still only lightly investigated part of the user-record layout -- [UNKNOWN] beyond "it's there and looks like a path".) - The fix: this is the same structure as Session 1 found, not a
different one -- only the case of the username differs. Confirmed by
recomputing
channel_table_start = offset_of("super") - 8 - chans_max*128with the lowercase match: 972012, exactly matching the independently-derived answer from the anchor-free scan. Checked the other two failing files (DB_EM_MountGordon_1003.gdb,DB_Mag_MountGordon_1003.gdb,DB_Mag_Elaine_1003.gdb) and all three also use lowercase"super"(all at the identical offset 139840, since all three sharechans_max=50), with zero occurrences of uppercase"SUPER"anywhere. Revision applied and verified:find_channel_table()now searches for both cases. All 9 files (2 USGS + 7 GSQ) now resolve correctly with the same underlying formula. - Interpretation: the two 2020 USGS files (Session 1) both happened to use uppercase; five of seven real 1990s-2020s GSQ files use lowercase. This looks like a difference in whatever tool/pipeline/Geosoft-version wrote each file's default superuser name, not a difference in the file format -- the byte-offset arithmetic, the 128-byte stride, and the "channel table immediately precedes the user table" structure all held up perfectly once the case was corrected. Logged as [CONFIRMED, with a documented revision] rather than silently patched.
- Second robustness issue found and fixed in the same investigation:
the literal word "SUPER"/"super" is not a fully reliable anchor even
when present -- confirmed by triggering a false-positive lock-on
during debugging (a stray occurrence of "SUPER" deep inside binary
data pointed at a nonsense offset before the fix in §2.4 below was
applied).
find_channel_table()now tries every occurrence in order and verifies the implied table start actually decodes a clean name before accepting it (see code comments ingdb_reader.py).
2.4 A second, independent real-file finding: unused channel-table capacity isn't always zeroed¶
- Observation:
DB_Mag_293.gdb(Holroy River, 1992) crashed the original reader withUnicodeEncodeErrorwhile printing channel names. Investigating record index 4 onward (ofchans_max=20, only 4 real channels:Easting_AMGz55, Northing_AMGz55, Mag, Rad_Alt) showed the "unused" slots contain non-zero, binary-float-looking leftover bytes, not the clean zero-padding seen in every 2020 USGS record.DB_Mag_833.gdbshowed the same pattern, and in that file two of the garbage slots happened to decode a clean-looking printable name ("L2161", a stray fragment of the projection-name blob from §2.3; and"1") that would otherwise have been accepted as bogus extra channels. - Fix, and how it was validated: added
_read_name()cleanliness detection (NUL-terminated + fully printable = clean; anything else = leftover garbage, not a real channel) and a secondlooks_sanecheck (real channel records always have aGS_*-valid or small-negative dtype code and a knownDB_CHAN_FORMAT_*code; the"L2161"/"1"garbage entries have wildly out-of-range values likedtype=2048orformat=-22912that immediately fail this check). Re-ran all 9 files after the fix: every one now reports exactly the number of real channels (spot-checked by eye against plausible survey channel lists), with zero false positives and zero crashes. - Takeaway for NOTES.md: "unused symbol-table capacity is always cleanly zeroed" was an assumption baked into Session 1's reader that turned out to be true only for the two 2020 USGS files and false for at least 2 of 7 real 1990s GSQ files. Recorded as a genuine, reader-breaking finding from the pressure test, now handled instead of assumed away.
2.5 The VA / array-channel question -- decoded and independently cross-validated¶
This was the main event of the session. Order of discovery:
- Round 1 (GEOTEM/Questem decay-gate channels): ran the (by-now
fixed) reader against
DB_EM_293.gdb,DB_EM_833.gdb,DB_EM_MountGordon_1003.gdb. All three list their per-gate decay data as N separate flat scalar channels:GEOTEMCh1..GEOTEMCh16+GEOTEM_50HZ_1/GEOTEM_50Hz(Holroy, 18 gate - housekeeping channels),
GEOTEMCHANNEL1..15(Scrubby Knob),CH1..CH15(Mount Gordon, Questem system -- different vendor, same one-scalar-channel-per-gate pattern). No VA/array channel found in any of these three independent real deliveries. This is a real, if partial, answer on its own: at least three independent real-world TEM processing pipelines (two different acquisition systems, three different years) chose the "N scalar channels" representation over a single array channel for per-gate decay data. - Round 2 (the actual VA channel, found): the operator's lead
pointed at
AG106386_Northern Georgetown_Conductivity.gdb(317MB, a modern GA/GSQ airborne EM conductivity delivery, structurally very different from the older files --chans_max=500,page_size=32768vs. the usual 1024). Ran the reader: 55 channels, all still showingGS_DOUBLE/GS_FLOAT/GS_LONGwith no visible array marker in the fields decoded so far -- but several channel names strongly implied array data (dBdt_X_Raw,dBdt_X,dBdt_X_TC_F, etc. -- a TEM system's raw/corrected/filtered decay curves, andLEI_Conductivity/LEI_Depth, a layered-earth-inversion depth model). - Test: dumped full raw 128-byte records for a definitely-scalar
channel (
Easting) side by side with the suspected-array channels and diffed them byte-by-byte. - Found it: a 2-byte int16 field at relative offset +118
reads
1for every plain scalar channel checked (Easting,Fiducial,LEI_Conductivity_0_5m, etc.) and a much larger value for the suspected array channels --24for everydBdt_X_*/dBdt_Z_*/Filter_Smoothness_*channel,30for bothLEI_ConductivityandLEI_Depth. - Systematic tabulation across all 55 channels confirmed this
holds with zero exceptions: every channel with
array_width==1is a genuinely scalar quantity; every channel witharray_width>1is one whose name independently implies multi-value data (decay gates, depth layers). This field is the answer to the "array-per-cell (VA) channel" question: relative offset +118 (int16) in the 128-byte channel record is the array width (elements per fiducial "cell");1= ordinary scalar channel (the overwhelming majority of real channels seen across all 10 real files this project has now examined),>1= a true VA/array channel. - Also noticed and tabulated a second field, relative +86 (int16),
whose value range (0, 1, 2 seen) matches the vendor's own
DB_ARRAY_BASETYPE_*enum (Session 1 §1.5:NONE=0, TIME_WINDOWS=1, TIMES=2, ...) --1(TIME_WINDOWS) appears on most (not all) of thedBdt_*/Filter_Smoothness_*array channels, which makes physical sense (decay-gate data keyed by time window). However this field is not a reliable array-ness indicator by itself:LEI_Conductivity/LEI_Depth(both real arrays, width 30) have+86=0; the three older GEOTEM/Questem files have+86=2(TIMES) on every single channel including obviously-scalar ones likeEasting/FID. Logged this field as [LIKELY] matchingDB_ARRAY_BASETYPE_*in name, but its exact write-time semantics (per-channel physical domain tag? sometimes a database-wide default instead?) are [UNKNOWN] -- flagged rather than overclaimed. The array-width field (+118) is the one doing the real work and is [CONFIRMED]. - Independent cross-validation via the public ASEG-GDF2 standard
(operator-directed lead, fetched and read myself rather than trusting
the description): downloaded the actual spec PDF from
https://www.aseg.org.au/public/200/files/ASEG-GDF2-REV4.pdf("THE ASEG-GDF2 STANDARD FOR POINT LOCATED DATA", Draft 4, ASEG Standards Committee, 27 Jan 2003) -- a genuine, unrelated, public Australian industry standard (Australian Society of Exploration Geophysicists) for plain-ASCII geophysical point data, with zero connection to Geosoft's binary internals. Appendix 1 ("THE ASEG-GDF2 SYNTAX") states the field-format grammar explicitly:nFw.d= "real in floating point form", where "n represents the number of repeats for array definitions, w represents the number of characters used, and d represents the number of decimal places." This is the exact, unambiguous, standards-body definition of the Fortran-style repeat-count syntax used inNorthern Georgetown_Conductivity.dfn(extracted fromNorthern Georgetown_Conductivity.zip, read directly, no parsing ambiguity):Both explicitly declared as 30-repeat fields, i.e. 30-element arrays -- matching the binaryDEFN 26 ST=RECD,RT=;LEI_Conductivity:30F10.4:NULL=-999.9999,UNIT=S/m,NAME=Conductivity derived with GA_LEI DEFN 27 ST=RECD,RT=;LEI_Depth:30F12.1:NULL=-99999999.9,UNIT=m,NAME=Depth to top of layer, derived with GA_LEI.gdb'sarray_width=30for these exact two channel names exactly. Every other field in the same.dfn(Easting, Northing, GA_project_number, the 10 individualLEI_Conductivity_*mchannels, etc.) has no repeat count (implicit 1), matchingarray_width=1for every one of those in the binary file. This is a complete, independent, standards-based confirmation of the array-width field -- one public source (binary structure, reverse engineered from real bytes) and one totally unrelated public source (a 25-year-old plain-text industry standard, read for its own stated purpose) agree exactly on which channels are arrays and how big. - Went one step further and confirmed actual row-level ASCII ground
truth matches the array structure, not just the channel-level width
claim: stream-read (via Python's
zipfile, without extracting the ~1GB file) the first few rows ofNorthern Georgetown_Conductivity.datstraight out of the zip. Confirmed by direct token-counting against the.dfnfield order that each real data row contains exactly: 25 scalar values (fields 1-25,GA_project_numberthroughPLM), then a run of exactly 30 conductivity values (LEI_Conductivity), then a run of exactly 30 monotonically increasing depth values (LEI_Depth,0.0, 3.0, 6.3, 9.9, 13.9, ... 445.9), then 10 more scalar values (the individualLEI_Conductivity_*mchannels) -- 95 tokens per row total, exactly matching the.dfn's declared field list. First row'sGA_project_number= 5027,NRG_Job_Number= 2347,Fiducial= 113120 (then 113140, 113160, 113180, ... incrementing by 20 on subsequent rows).
2.6 The third storage/compression scheme -- found and structurally characterized¶
Prompted by an explicit reminder not to let the documented-but-unobserved
third scheme (DB_COMP_SPEED/DB_COMP_SIZE, both zlib-based per Session
1 §1.7A/S5 -- Session 1 had only observed DB_COMP_NONE in real files)
quietly drop out of the documentation, and a follow-up lead from the
operator pointing specifically at AG106386_Northern
Georgetown_Conductivity.gdb as a likely real example.
- First signal: this file's header has
page_size = 32768at the usual offset 100 -- the only real file examined so far (10 total now) with anything other than the default/documented-normal 1024. A much larger page size is a reasonable prior for "this database is paged for compression" (bigger pages amortize per-page compression overhead better). - Test: scanned the file's data region (past the symbol table) for
the standard zlib stream header byte pairs (
78 9c,78 01,78 da,78 5e-- the four common zlib compression-level/window combinations). Found dozens of hits for each. Crucially, the78 01hits are exactlypage_size(32768) bytes apart: 6248, 39016, 71784, 104552, 137320, ... (differences all exactly 32768). This is not a coincidence -- it means every single page-sized slot in this region of the file begins with its own independent zlib stream. - Verified by actually decompressing:
zlib.decompress()on the bytes starting at one of these offsets succeeds cleanly. Usingzlib.decompressobj()(which reports unconsumed trailing bytes) showed the real compressed stream inside one 32768-byte page slot was only 234 bytes long, decompressing to 42216 bytes of a single int32 value (2347) repeated 10554 times -- i.e. one entire channel's worth of data for a constant column, which explains the huge (~180x) compression ratio trivially. The remaining ~32.5KB of the page slot is padding; almost all zero, except a small nonzero region right near the very end of the slot (~72 bytes before the boundary) that looks like a short per-page trailer/footer (contains what appears to be the decompressed size,0xa4e8=42216, matching exactly) -- [UNKNOWN] beyond that; not fully decoded, flagged for future work. - This directly answers the operator's core ask: confirmed a real
.gdbfile using genuine zlib-based page compression, distinct in layout fromDB_COMP_NONE(Session 1's two USGS samples, whose data is stored as flat uncompressed column runs findable by direct byte search) -- and structurally different from the.grdsibling format's compression scheme (Session 1 §1.9/§4:.grduses an explicit offset-table + size-table pointing at variable-length compressed blocks with a 16-byte per-block sub-header; this.gdbfile instead uses simple fixed-page_size-stride placement with no separate index table found so far -- each page slot is a constant 32768 bytes regardless of how much of it the actual compressed stream uses). - Found the compression-level header field: compared the full
header int32 dump (offsets 0-124) between a known-
DB_COMP_NONEfile (Magnetic_Data.gdb) and this now-confirmed-compressed file. Offset 120 is 0 in the uncompressed file and 2 in the compressed one -- matchingDB_COMP_NONE=0/DB_COMP_SIZE=2from the vendor's own enum (Session 1 §1.5) exactly. [LIKELY] (single positive + single negative real-file example is a real but not exhaustive test; promoted from "unconfirmed header offset" to "likely comp-level field").DB_COMP_SPEED=1has still not been observed in any real sample -- explicitly left as[UNKNOWN -- not yet observed]per the operator's instruction not to let it quietly disappear. - The capstone confirmation -- exact numeric ground truth, not just
structure: having independently obtained real ASCII row values for
this exact survey via the ASEG-GDF2
.datfile (§2.5), searched for and found the actual compressed pages corresponding to the first three channels in symbol-table order, right where column-major ordering (Session 1's other big confirmed finding) predicts them to be: - Page at file offset 3,473,480 decompresses to the constant int32
value 5027, repeated -- exactly the real
GA_project_numbervalue from row 1 of the real.datfile (channel index 0). - Page at file offset 3,506,248 (exactly
+32768from the previous) decompresses to the constant int32 value 2347, repeated -- exactly the realNRG_Job_Numbervalue from row 1 (channel index 1). - Page at file offset 3,539,016 (another
+32768) decompresses to 113120, 113140, 113160, 113180, 113200, 113220, ... -- exactly the realFiducialvalues from rows 1, 2, 3, 4, 5, 6 of the.datfile, in order, incrementing by 20 each row exactly as the real ASCII data does (channel index 2).
This is a complete, end-to-end, value-exact validation chain: public binary structure (page-aligned zlib) → decompressed by nothing but the Python standard library → real integers → matched digit-for-digit against an independently-sourced, standards-defined, plain-text export of the same real survey. Nothing about this result depends on Geosoft's engine at any point. This is the strongest single result of Session 2, arguably of the whole project so far. - Still open: the exact offset/index structure that lets a reader locate the first data page and know how many pages belong to each channel without brute-force scanning for zlib magic bytes (the approach used here, which works but isn't how a real reader should operate). Header offset 104 ([UNKNOWN] in Session 1) is a plausible candidate -- it reads 3,408,620 in this file, reasonably close to (but not exactly matching) where the real compressed data was found to begin (~3,473,480) -- logged as a lead, not a confirmed answer.
2.7 Reader updated and re-validated end to end¶
reader/gdb_reader.py updated to reflect every finding above:
case-insensitive super-user search with false-positive rejection,
garbage-slot detection (name_is_clean) and record sanity-checking
(looks_sane), array-width (array_width/is_array) and array-basetype
decoding, a comp_level header field, and a relaxed magic check (4-byte
hard requirement, 16-byte common-case reported as a note not an error).
Re-ran against all 10 real files now in the project (2 USGS + 7 GSQ
EM/Mag + 1 GSQ AG106386) after every change; all 10 parse cleanly with
plausible channel lists and no crashes as of the final version.
2.8 Session wrap¶
NOTES.md needs a substantial update: revise the SUPER-anchor claim to
note the case-sensitivity fix, add the unused-capacity-not-always-zeroed
finding, add the full VA/array-channel section (this is now one of the
best-evidenced findings in the whole project), add the third
compression-scheme section, and update the sample-file table with the
GSQ files and the ASEG-GDF2 spec as a new cited source. Reader and its
docstrings already updated in place as the source of truth for the exact
byte offsets; NOTES.md should summarize and cross-reference rather than
duplicate at full length.
2.9 CORRECTION: DB_COMP_SPEED is not actually unobserved -- re-investigation¶
After writing §2.6/§2.8 and reporting "DB_COMP_SPEED=1 not yet observed
in any real file" to the operator, the operator pushed back, noting that
several of the other new GSQ files (not the Georgetown conductivity
one) appeared to report a container-level compression type of "Speed."
Provenance note, for accuracy: this was not a finding from separate
tooling or an external method -- the operator looked at the same header
offset (120) this project had already identified and was already using
to tell DB_COMP_NONE from DB_COMP_SIZE, and noticed it held the third
documented value on files this project hadn't gotten around to checking
yet. So the "independent check" was this project's own field, applied
more broadly than this project itself had -- not a new technique or an
outside source. This section is the honest re-investigation of that
discrepancy, and it was a real miss -- the fix is straightforward in
hindsight: §2.6 only ever checked header offset 120 on AG106386 vs.
the two USGS files. It was never re-checked on the 7 GEOTEM/Questem GSQ
files from earlier in the same session. Re-checking that one field
immediately resolved the question:
| File | Offset 120 | Offset 100 (page_size) |
|---|---|---|
DB_EM_293.gdb |
1 | 1024 |
DB_Mag_293.gdb |
1 | 1024 |
DB_EM_833.gdb |
1 | 1024 |
DB_Mag_833.gdb |
1 | 1024 |
DB_EM_MountGordon_1003.gdb |
0 | 1024 |
DB_Mag_Elaine_1003.gdb |
0 | 1024 |
DB_Mag_MountGordon_1003.gdb |
0 | 1024 |
AG106386_...gdb |
2 | 32768 |
| Both USGS files | 0 | 1024 |
Correction: DB_COMP_SPEED=1 is observed, in 4 real files (both
.gdb files from each of the Holroy River and Scrubby Knob deliveries).
The Mount Gordon files (all 3) and both USGS files are DB_COMP_NONE=0.
Only AG106386 is DB_COMP_SIZE=2. LOG.md §2.6/§2.8 and the earlier
report to the operator were simply wrong on this point because the field
wasn't checked broadly enough -- corrected here, not glossed over.
Given the field is real and present, tracked down what it actually
means for the data -- per the operator's explicit steer not to assume
Speed uses the same algorithm as Size just because Size turned out to be
zlib, and to apply the same "don't trust the container's own claim"
skepticism here that the Loop3D .grd finding already taught (COMP_TYPE
said LZRW1, bytes were actually zlib -- so a claim in either direction
needs a real byte-level check, not an assumption).
- First (wrong) assumption, caught and corrected in real time: on
first pass, scanning
DB_EM_293.gdbfor the zlib magic byte pair78 01found 1771 hits, with 1759 of them (99.3%) clustering at a constant residue (byte 63) modulo the 1024-byte page size -- overwhelming evidence of a deliberate structural pattern, not noise. Assumed this meant "just likeAG106386, this is a zlib stream starting at a fixed page offset" and triedzlib.decompress()there. It failed (invalid stored block lengths/incorrect header checkdepending on exact offset tried) -- the first sign something was different about Speed mode, not simply "zlib at a different offset." - Root cause of the false lead: dumping the full bytes around one of
these hits revealed the real structure:
... 0f 0e ff fe 12 34 56 78 01 00 00 00 00 00 00 00 ...-- the same 16-byte magic sub-header already found in two other places this project (the.grdsibling format's per-block header, Session 1 §4; andAG106386's confirmed zlib pages, §2.6) --0f 0e ff fe, then the literal constant12 34 56 78, then two more int32 fields. The "78 01" byte pair my search had locked onto was never a zlib header at all -- it was just the tail end of this magic constant (...56 78) immediately followed by the low byte of the next field (01 00 00 00), a coincidental substring match, not a real signal. Lesson re-learned the hard way: a byte-pair match is not evidence on its own without checking what actually surrounds it. - The real, confirmed structural finding, once this was sorted out:
searched for the actual 8-byte magic (
0f 0e ff fe 12 34 56 78) directly instead of the misleading 2-byte substring. Found it 1761, 320, 2160, and 600 times respectively in the fourDB_COMP_SPEED=1files, and -- checking the int32 field immediately following the magic (relative +8 within the 16-byte header) -- it reads exactly1in effectively 100% of instances (1759/1761, 320/320, 2160/2160, 600/600 -- the handful of exceptions in the first file look like coincidental false-positive magic matches inside actual data, not a real second subtype). Ran the identical search againstAG106386(confirmedDB_COMP_SIZE=2): the same magic appears 3414 times, and the field reads exactly2in all 3414 instances. This field is a per-page/per-block compression-subtype tag, and it matchesDB_COMP_SPEED=1/DB_COMP_SIZE=2from the vendor's own enum (§1.5) exactly, with zero exceptions across two files and thousands of real instances. This means the 16-byte magic wrapper is a genuinely shared low-level container primitive used across (at least) three contexts now:.grdcompressed blocks,.gdbDB_COMP_SIZEpages, and.gdbDB_COMP_SPEEDpages -- with this one field distinguishing which specific compression algorithm is used inside. - Checked whether the Speed payload is zlib -- directly, skeptically,
the same way Size was checked, per the operator's explicit
instruction, rather than assuming the vendor doc's "both tiers use
zlib" claim just because it turned out true for Size: the byte
immediately following the 16-byte header in a confirmed
AG106386(Size) page is78 01-- a valid zlib CMF/FLG pair, andzlib.decompress()on it succeeds (already established, §2.6). The byte immediately following the identical 16-byte header in a confirmed Speed page (DB_EM_293.gdb) is23 00 00 92 07 00 00 ...-- not a valid zlib header (0x78is required as the first byte for the standard 32K-window deflate stream;0x23isn't a valid zlib CMF value), andzlib.decompress()/decompressobj()fail immediately at every byte offset tried in the neighborhood (0 through +40 past the header). Result:DB_COMP_SPEED's payload is confirmed NOT to be a standard zlib stream, directly contradicting Seequent's own documentation (S5), which states both compression tiers use "the lossless open source library at zlib.net." This is a genuine, verified vendor-documentation discrepancy -- not on the sibling.grdformat's internalCOMP_TYPEfield this time (which is a different kind of claim, an on-disk enum), but on Seequent's own published end-user help documentation for.gdbitself. - Tested the LZRW1 hypothesis directly, since it's a real, named,
publicly-documented algorithm this project already had a reason to
suspect (the task's own source (e); LZRW1's known performance profile
-- very fast, modest compression ratio -- matches the qualitative
difference Seequent's docs describe between the two tiers: "speed"
mode is faster but compresses less (58% of original size) than "size"
mode (19%), which is exactly the kind of trade-off LZRW1 (fast,
lighter compression) vs. zlib/deflate (slower, better ratio) would
produce. Fetched Ross Williams' own canonical public-domain LZRW1
reference implementation directly (
http://www.ross.net/compression/download/original/old_lzrw1.c-- the exact C source from his 1991 Data Compression Conference paper, explicitly marked public domain in its own header comment), ported its decompression routine (lzrw1_decompress) to Python by hand (kept inscripts/lzrw1_decode.pyand a no-flag-prefix variant inscripts/lzrw1_decode2.py), and tried it against the real Speed-mode payload bytes at every candidate starting offset from the end of the 16-byte magic header out to +40 bytes (accounting for the reference implementation's own 4-byteFLAG_BYTESprefix convention, and a variant without it, in case Geosoft's implementation omits that detail). Result: inconclusive. No candidate offset produced a long, clearly-valid, sensible-looking decoded output the way the correct zlib offset did for Size mode on the first reasoned attempt. A couple of offsets ran to completion without an internal self-reference error and produced output of a plausible length, but the decoded bytes (checked as both raw hex and as candidate float64 values) didn't show an obviously-real, unambiguous pattern the way the Size-mode/AG106386 result did (constant repeated integers and incrementing fiducials, matching independent ground truth exactly). - Honest conclusion, per the operator's own framing of what an
acceptable outcome looks like here:
DB_COMP_SPEEDis real, present in 4 real files, and demonstrably uses the same low-level container wrapper (the 16-byte magic header) asDB_COMP_SIZEand.grd's own compression scheme, with a correctly-matching subtype tag. But its actual payload compression algorithm is not standard zlib (verified directly, not assumed) -- contradicting Seequent's own published documentation, which is itself a legitimate and useful finding, not a dead end. Whether it's genuinely LZRW1, a related LZRW-family variant (LZRW1-A, LZRW2, LZRW3, etc. -- Ross Williams published several), or something else entirely (a Geosoft in-house scheme) was not resolved with confidence in the time available. This is recorded inNOTES.mdas an open problem with the same rigor as everything else -- what's confirmed (the field is real, the wrapper is shared, it's not zlib) is separated clearly from what's still unknown (which algorithm it actually is).
2.10 BREAKTHROUGH: DB_COMP_SPEED fully decoded -- it is canonical LZRW1, exactly¶
Following a further operator steer: don't give up on LZRW1 just because the exact reference-C item encoding didn't validate at a byte-offset search -- test the more basic, more likely-to-survive-a-customization claim first (the group-of-up-to-16-items-per-control-word framing itself, independent of assuming any particular item byte width), using real backreference validation rather than plausibility-by-eye. This directly found the answer.
Method: wrote a real, validated decoder (not just a byte-counting scan) parameterized by (start-offset-past-the-16-byte-magic-header, literal-item-width). "Validated" means: for a hypothesized parameter set, actually walk control-word groups and REQUIRE every copy-item's backreference to point at an already-produced position in the (still hypothetical) output -- exactly the correctness constraint a real decoder must satisfy, as opposed to just checking there are "enough bytes left" (which is nearly meaningless on its own -- an earlier, cruder version of this test using only byte-accounting without backreference validation gave misleadingly strong-looking "survival" results for large literal widths that turned out to be an artifact of the weak check, not a real signal; caught and discarded before drawing any conclusion from it).
Ran this validated decoder across (start_offset in 0..15) x
(literal_width in {1,2,4,8}) against 200 independent real chunks from
DB_EM_293.gdb, counting how many consecutive groups each parameter
combination survives before hitting an invalid backreference. Result was
completely unambiguous: literal_width=1 (i.e. plain single-byte
literals, exactly like canonical LZRW1) at start_offset=12 survived
200/200 chunks for at least 10 groups (average 249 groups, max 703),
while every other parameter combination in the entire 64-combination
grid survived on average ~1 group (i.e. failed almost immediately) --
not a close contest, a total outlier.
Followed up with an exact-length decode (stop precisely at the
declared/assumed output length rather than an approximate cap) and found
the real per-chunk framing exactly: past the already-known 16-byte magic
sub-header (0f 0e ff fe 12 34 56 78 <subtype=1> <reserved>), there is
a 12-byte length sub-header: <decompressed_length int32>
<chunk_length int32> <marker int32>, where chunk_length includes
these 12 bytes (chunk_length - 12 == the number of raw compressed
bytes that follow), and marker is a constant, 0xF4E5D6C7
(-186263865 signed), confirmed identical across all four real Speed
files, effectively 100% of the time (2000+ chunks checked, one stray
exception attributable to a coincidental magic-byte false-positive
elsewhere in real data, same class of noise seen elsewhere in this
project).
Decoding exactly decompressed_length bytes of plain canonical
LZRW1 (Ross Williams' reference lzrw1_decompress() core loop,
ported directly to Python -- 2-byte control word, 1-byte literal items,
2-byte nibble-packed copy items, no 4-byte FLAG_BYTES prefix from
his C wrapper) starting right after that 12-byte header consumes
exactly chunk_length - 12 input bytes, with zero slack, in
every one of 30/30 chunks sampled from each of the four real Speed files
(DB_EM_293.gdb, DB_Mag_293.gdb, DB_EM_833.gdb, DB_Mag_833.gdb;
120/120 total). This is not a plausibility argument -- it's an exact,
falsifiable, repeatedly-passed numeric check: the encoder's own recorded
compressed length matches what an independent, from-scratch canonical
LZRW1 decoder actually consumes, to the byte, every single time.
The decoded bytes are also independently sensible, not just
byte-count-consistent: decoding several consecutive real chunks from
DB_EM_293.gdb and interpreting the output as float64 produces smoothly
varying, physically plausible-looking values (e.g. one early channel's
chunks decode to sequences like 10764.0, 10194.0, [dummy], 10319.0,
9947.0, [dummy], 9465.0, 9141.0, ..., smoothly decreasing across
successive chunks in a pattern very consistent with real TEM
decay-curve-like data), and the dummy/no-data sentinel value exactly
matches rDUMMY = -1.0E32 from the vendor's own published constants
(Session 1 section 1.5) wherever it appears in the decoded stream.
Checked whether the same 12-byte length sub-header applies to
DB_COMP_SIZE (zlib) chunks too, for completeness/symmetry: it does
not -- confirmed the zlib payload in AG106386 starts immediately at
magic_offset + 16 with no 12-byte gap (already established in
sections 2.6/2.9; re-confirmed here by testing the 12-byte-header
hypothesis against Size-mode chunks and finding it produces nonsense
decompressed_length values, as expected once you know there's no such
header there). This is a genuine, confirmed structural asymmetry between
the two compression modes' chunk framing, not an oversight: Speed-mode
chunks carry their own explicit length pair; Size-mode chunks apparently
don't need to (zlib streams are self-terminating, so no length prefix is
strictly required to decode them, only to know where the next chunk
starts without decoding -- which may itself explain why Speed mode
needs the extra 12 bytes and Size mode doesn't, if Size-mode readers are
expected to locate chunk boundaries some other way, e.g. via the
still-unidentified line/channel index, section 6.4/8 of NOTES.md).
Final, complete, high-confidence result: DB_COMP_SPEED is Ross
Williams' canonical LZRW1 algorithm, byte-for-byte, wrapped in a
Geosoft-specific 28-byte chunk header (16-byte shared magic + 12-byte
length/marker fields). This also resolves the vendor-documentation
discrepancy noted in section 2.9 more precisely: Seequent's help
documentation states both compression tiers use zlib; that is simply
incorrect for the Speed tier, which uses LZRW1 -- a different, older,
faster, and (per LZRW1's well-known profile) lower-ratio algorithm, that
just so happens to be exactly the algorithm the task's own source (e)
anticipated might be relevant somewhere in this format family, even
though it turned out NOT to be the answer for the sibling .grd format
or for .gdb's Size mode (both zlib, confirmed). LZRW1 really is in
here -- just not where it was originally guessed.
Working code: reader/lzrw1.py (clean, final, documented decoder + chunk
finder, tested against all 4 real Speed files). Exploratory/search
scripts kept for provenance in scripts/lzrw1_decode.py,
scripts/lzrw1_decode2.py, scripts/lzrw1_group_search.py,
scripts/lzrw1_element_decode.py (this last one tests a different,
now-refuted hypothesis -- 8-byte-element literals -- kept because a
refuted hypothesis with its negative result is still part of the honest
record, not just the ideas that worked).
2.11 Operator asked: is "all 4 real Speed files" actually all of them? Re-check + full non-sampled validation¶
The operator asked to re-check header_fields()['comp_level'] across
every real .gdb file sitting in samples/GSQ_Data -- including
ones extracted for earlier rounds but never checked for this
(Melinda-Downs-1.zip, Melinda-Downs-2.zip), and the two large,
previously-skipped archives (Kamilaroi.zip, Georgetown-AGSO.zip)
-- rather than trusting "the four files I happen to remember checking."
This was a completely fair challenge to the §6.5c/§2.10 result's actual
scope, and the answer was: no, "all 4" was not actually all of them.
Re-check, properly this time:
- Extracted Melinda-Downs-1.zip -> gg001213/ (DB_AGG_1213.gdb,
DB_Mag_1213.gdb) and Melinda-Downs-2.zip -> gg001212/
(DB_AGG_1212.gdb, DB_Mag_1212.gdb) -- never previously extracted
or looked at in this project at all. AGG = airborne gravity
gradiometry, a data type not previously represented in any sample.
- For the two large, deliberately-skipped archives, did not fully
extract them (multi-GB, impractical) -- instead used Python's
zipfile module to open individual entries as streams and read just
their first 128 bytes, which works without extracting the whole
archive (or even the whole entry) to disk:
- Kamilaroi.zip: Mag_100100.gdb, Rad_100100.gdb,
DEM_100100.gdb -- all three comp_level=0 (DB_COMP_NONE).
Nothing new here; not extracted further.
- Georgetown-AGSO.zip: DB_Rad_1027.gdb and DB_Mag_1027.gdb are
both comp_level=1 (Speed -- previously unchecked!);
DB_DEM_1027.gdb is comp_level=0. Extracted just the two Speed
files from the archive (108MB and 807MB) using unzip with
explicit entry names, rather than extracting the full 1.5GB zip.
- Running header_fields()['comp_level'] across literally every
.gdb file now reachable in samples/GSQ_Data (14 files: 8 already
known + these 6 new ones) found 10 real DB_COMP_SPEED files total,
not 4: the original DB_EM_293.gdb, DB_Mag_293.gdb,
DB_EM_833.gdb, DB_Mag_833.gdb, plus 6 more:
DB_AGG_1213.gdb, DB_Mag_1213.gdb, DB_AGG_1212.gdb,
DB_Mag_1212.gdb, DB_Rad_1027.gdb, DB_Mag_1027.gdb. This is a
second, larger instance of the exact same class of mistake as §2.9 --
scope-checking a claim only against the files already in hand rather
than every file actually available -- and worth calling out plainly
rather than downplaying: the original "[CONFIRMED] against all 4 real
Speed files" claim understated the real sample by more than half.
Full, non-sampled validation (every chunk, not the first N) against all 10 files -- and a second real, honestly-investigated finding, not just a clean sweep:
Running the full validation (script: scripts/lzrw1_full_validation.py)
against the original 4 files reproduced the earlier 100% result exactly
(1759/1759, 320/320, 2160/2160, 600/600 chunks, zero failures). Running
it against the 4 new Melinda Downs files (DB_AGG_1213.gdb,
DB_Mag_1213.gdb, DB_AGG_1212.gdb, DB_Mag_1212.gdb) initially
showed real failures -- 209/640, 49/640, 145/448, and 28/448 chunks
respectively failed to decompress (invalid backreference errors) with
the decoder as it stood after §2.10. Rather than treating this as "the
algorithm sometimes doesn't work" and moving on, investigated why:
- Noticed the failure count exactly equalled a
marker_mismatchescount in every file (i.e. every chunk with amarkervalue other than the known constant0xF4E5D6C7also failed to decode, and every chunk that decoded successfully had that exact marker) -- too clean a correlation to be coincidental noise. - Dumped the raw bytes of one failing chunk's header and found its
markerfield reads a different, specific, non-random constant:0xF0E1D2C3. Compared byte-by-byte against the known-good0xF4E5D6C7: every byte differs by exactly0x04(f4-f0=4,e5-e1=4,d6-d2=4,c7-c3=4) -- clearly a deliberate, related second constant, not noise. - Checked the failing chunk's own recorded lengths:
chunk_length - 12equalsdecompressed_lengthexactly for every failing chunk (i.e. zero compression ratio) -- and readingdecompressed_lengthbytes directly, with no decompression at all, produces smooth, physically plausible float64 values (checked: real-looking gravity readings, e.g.88.08, 87.56, 86.47, 86.11, ...). This is exactly Ross Williams' own reference implementation'sFLAG_COPYcase (used when LZRW1 compression doesn't shrink a block, so the encoder just stores it raw instead of wasting compressed-format overhead on incompressible data) -- Geosoft's on-disk variant apparently repurposes the 12-byte header'smarkerfield itself to signal this, rather than a separate flag byte the way Ross Williams' C wrapper does. - Verified this explains every single failure with zero exceptions:
re-scanned all 4 Melinda Downs files' chunks and confirmed each one's
marker is one of exactly two values (
0xF4E5D6C7= compressed,0xF0E1D2C3= stored raw) with zero third values, and every stored-raw chunk satisfieschunk_length - 12 == decompressed_lengthexactly (209/209, 49/49, 145/145, 28/28). - Updated
reader/lzrw1.pyto handle both cases (SpeedChunk.is_stored_raw/.is_compressed,MARKER_COMPRESSED/MARKER_STORED_RAWconstants,decode_speed_chunk()branches on marker). Re-ran the full, non-sampled validation against all 8 chunk-scheme files: 100% success, zero failures, zero unrecognized markers, every single chunk in every file (DB_EM_293.gdb1759/1759,DB_Mag_293.gdb320/320,DB_EM_833.gdb2160/2160,DB_Mag_833.gdb600/600,DB_AGG_1213.gdb640/640 [431 compressed + 209 stored-raw],DB_Mag_1213.gdb640/640 [591+49],DB_AGG_1212.gdb448/448 [303+145],DB_Mag_1212.gdb448/448 [420+28]).
A third, separate, genuinely new finding from the last 2 of the 10
files (DB_Rad_1027.gdb, DB_Mag_1027.gdb, both from
Georgetown-AGSO.zip): running the same chunk scan on these found
zero and two magic-header hits respectively -- essentially none,
and the two found in DB_Rad_1027.gdb are both subtype=2
(DB_COMP_SIZE-tagged, not Speed), almost certainly coincidental noise
given how rare real hits are relative to file size elsewhere. Neither
file uses the chunked scheme at all, despite both declaring
comp_level=1 (DB_COMP_SPEED) in their header. Checked whether their
actual channel data is stored some other compressed way, or plain
raw: searching for the longest run of "plausible-looking" float64
values (same technique as Session 1, §1.16) found very long clean runs
in both files -- 1628 consecutive plausible values in DB_Rad_1027.gdb
(smoothly-decreasing latitude-like values starting -18.00017,
-18.000788, -18.001403, ...) and 16314 consecutive plausible values in
DB_Mag_1027.gdb (similarly smooth, -18.000008, -18.000074,
-18.00014, ...). Both files' actual channel data is stored
completely uncompressed -- directly byte-searchable raw doubles, the
same as a DB_COMP_NONE file, not merely "mostly stored-raw chunks"
like the borderline cases in the Melinda Downs files, but with no
chunk-header machinery anywhere in the file at all.
Honest characterization of this last finding: the database-level
comp_level header field (offset 120) reflects the compression mode
the database was configured with, not a hard guarantee that any
particular byte of channel data was actually compressed with it. Two
real files nominally configured for DB_COMP_SPEED contain zero
compressed (or even chunk-wrapped-but-stored-raw) data -- the entire
file's data is laid out exactly like an uncompressed database. Both of
these are also, notably, much larger than any of the other 8 real Speed
files (108MB and 807MB vs. 2-12MB) -- worth flagging as a possible
(unconfirmed) correlation for anyone continuing this work: maybe there's
a size threshold or per-database (rather than per-chunk) decision point
involved, or maybe it's specific to this delivery/tool version. Recorded
as an open observation, not a resolved mechanism.
Bottom line, corrected and now fully scoped: DB_COMP_SPEED (=
canonical LZRW1, wrapped in the 28-byte chunk header with two possible
marker values) is fully validated against 8 of the 10 real files
that declare it, covering every single chunk in each (thousands of
chunks, zero failures). The remaining 2 files declare DB_COMP_SPEED
but contain no compressed data to validate against at all -- a real,
separate, and equally honestly-reported finding about what the
comp_level field does and doesn't guarantee.
Session 3 — the blob index (2026-09-09)¶
3.0 Starting conditions and the surviving lead¶
Resumed after a prior session was killed by an unrelated API-level
error mid-investigation of the single biggest open item flagged
throughout NOTES.md: the structure connecting a (line, channel) pair
to its data offset in the file. That session did not get to write
anything down. All that survived, relayed secondhand: "a real per-blob
header with the row count, GS type, and a blob index that matches the
channel's symbol-table index... [need to find] the master index table
that maps blob index -> file offset." Treated explicitly as an
unverified lead, not a fact, per the task brief.
Before doing anything else, checked the repo for any other trace of
that session's work. Found one: scripts/locate_channel_data.py,
present on disk but git-untracked (confirmed via git status) and
never mentioned in LOG.md/NOTES.md — consistent with being
mid-flight work from the killed session. Read it in full: it's a
never-completed harness (finds real CSV ground-truth values as raw
float64 bytes in the file, cross-references against
gdb_reader.read_channels() to report each match's symbol-table
index, sorted by offset) — a reasonable starting point for the
investigation described in the lead, but it doesn't yet implement or
demonstrate the blob-header/index finding itself. Did not delete it;
left in place and eventually formalized/extended its idea in
scripts/find_blob_index.py (§3.6).
3.1 Testing the lead directly against real bytes¶
Went straight to the byte offsets LOG.md/NOTES.md §1.16/§6.4
already had on record for Magnetic_Data.gdb's real, ground-truth-
verified column starts (fid=698416, raw_mag=721968,
comp_mag=745520, base=769072 — each real value's location, already
confirmed against the real CSV export in Session 1). Dumped 64 bytes
immediately before each of these four offsets.
Found a real, constant 48-byte structure immediately preceding every
one, always starting with the same 4-byte value CC CC 00 FF. Manual
field-by-field decode (all little-endian):
+0 (4 bytes) magic: CC CC 00 FF
+4 int32 n_pages_a
+8 int32 n_pages_b (always == n_pages_a in every real record seen)
+12 int32 "idx"
+16 int32 (looked like a plausible Unix timestamp)
+20 int32 200 (constant across all four)
+24 8 bytes zero
+32 float64 1.0
+40 int32 (varying-looking value)
+44 int32 5 (matches GS_DOUBLE -- and all four real channels here
ARE GS_DOUBLE per the symbol table)
+48 data starts here
The +12 field read 9, 13, 14, 15 for fid, raw_mag, comp_mag, base
respectively. Cross-checked immediately against
gdb_reader.read_channels()'s reported .index for each of those four
channel names: exact match, all four. This is precisely the part
of the surviving lead that was directly falsifiable, and it passed
immediately, first try. The +44 field (5) also matches each
channel's own symbol-table GS_* dtype code exactly. Both halves of
the lead ("row count... GS type... blob index matching the channel's
symbol table index") independently confirmed as real and correctly
described, on the very first real test.
3.2 The "row count" field, tested against real decoded values¶
The +40 field read 2841 for all four channels' first blocks.
Decoded exactly 2841 float64 values starting at each channel's known
data offset: fid's first 5 values are 577342.0, 577343.0, 577344.0,
577345.0, 577346.0 — the first value matches the real CSV ground
truth (fid=577342 from row 1) exactly, and the full run is a smooth,
monotonic-looking fiducial sequence, not garbage. Confirmed the +40
field really is a real, correct row count for this specific blob,
not a guess. Checked the leftover bytes after 2841*8 real data bytes
and before the next channel's 48-byte header (776 bytes): all zero —
consistent with each blob being allocated a fixed page-rounded
capacity with real data at the front and zero padding trailing.
3.3 The block stride is exactly n_pages * page_size — and the "idx" field decodes with one more insight than the lead had¶
The byte distance from one channel's header start to the next
(698368 -> 721920, etc.) is exactly 23552 bytes, every time.
23552 / 1024 (page_size) = 23 exactly — matching the +4/+8 fields
read from the very same headers (23, 23). This nails down n_pages
as literally "this blob's total allocated size, in pages" and confirms
the stride isn't a separate coincidence — it's computed from the
header's own n_pages field.
Broadened the marker search (data.find(b"\xcc\xcc\x00\xff", ...),
scanning the whole file rather than just the four known positions) to
see the idx field's full range. Found 1501 real, structurally-sane
matches (every single marker hit decoded a plausible header — zero
false positives) in the first 60MB of Magnetic_Data.gdb alone,
falling into repeating groups: {0,1,4,7,8,9,13,14,15}, then
{50,51,54,57,58,59,63,64,65}, then {100,101,104,107,108,109,113,
114,115}, then {150,151,154,...}, {200,...}, {250,...},
{300,...} — each group's members are the same set of channel
offsets ({0,1,4,7,8,9,13,14,15}) plus a constant that increases by
exactly 50 (== chans_max for this file) each time.
This is the actual master-index formula the killed session's lead was reaching for, one level more specific than what survived:
The surviving one-sentence lead ("a blob index that matches the channel's symbol-table index") is correct only for line slot 0 (whereblob_index == channel_slot_index trivially, since the line term is
zero) — which is presumably exactly what the killed session had been
looking at when it ran out of time, and a perfectly reasonable
observation as far as it went. The multiplication by chans_max is
the missing piece that makes it a general formula rather than a
coincidence specific to the first line.
3.4 Confirming line_slot_index is a real, physical line-table slot number — not just an arithmetic artifact¶
This needed independent confirmation, since dividing an integer by
chans_max will always "work" arithmetically regardless of whether
the quotient means anything. Used the real line table already on
record (LOG.md §1.16: line name "L1000" found at absolute offset
444320, +32 relative to record start, 128-byte stride, same table
noted in NOTES.md §6.3).
First attempt at walking backward from "L1000"'s record to find the
line table's true start had a real bug (read_line_name() treated an
empty/all-zero name field as cleanly empty rather than unused
capacity, since an empty string trivially passes an "all characters
printable" check on zero characters) — walked 701 records back through
what turned out to be legitimately-unused table capacity before
stopping on unrelated data, which would have implied "L1000" sits at
physical slot 701, contradicting the blob-index evidence (which implies
slot 0). Caught this by directly dumping the neighboring records
without the buggy filter: slot -1 (444160) has an all-zero name and
a 65536 value at the record's category field (+108, same field
used for line categories in NOTES.md §6.3); slot 0 (444288, "L1000")
has category 100 — matching DB_CATEGORY_LINE_NORMAL exactly; slot 1
(444416) is "L1001", category 100; slots 2-7 are "L1010",
"L1020", "L1030", "L1040", "L1050", "L1060", all category 100.
Confirmed: "L1000" really is physical slot 0 of the line table,
settled by a real structural marker (the category sentinel value),
not by the buggy scan. The earlier 701-record overrun was a real bug,
caught and explained rather than quietly patched around.
3.5 The capstone: two independent channels, one real line, both matching independent ground truth¶
With line_slot_index=0 confirmed to mean the real line "L1000",
computed blob_index = 0*50 + 10 = 10 for the line channel itself
(channel_slot_index=10 per read_channels() — this channel is a
64-byte string type, -64 dtype code, per NOTES.md §6.2). Walked the
file (via file seeks, not a fixed in-memory window, since this blob
turned out to sit at absolute offset 362,611,712 — far past the
64-channel window used for the numeric probe) until finding a header
with idx==10. Found it, decoded its declared row count (2841 — same
row count as the four numeric channels for the same line, exactly as
expected since all channels on one line share one fiducial range) as
fixed 64-byte NUL-padded strings.
Result: 2841 repetitions of the literal string "L1000". Matches
the real ground-truth line name exactly, for a completely different
channel (string-typed, not numeric) than the ones used to derive the
formula, stored in a completely different region of the file
(offset ~362MB vs. ~700KB for the numeric channels) — yet addressed by
the exact same formula and the exact same 48-byte header shape. This
is the strongest single confirmation obtained this session.
3.6 Removing the last "brute-force scan" dependency: finding where the chain starts, and walking it whole¶
The formula tells you which blob you want; it doesn't yet tell you
where the first blob is without a scan. Checked NOTES.md's existing
"live lead" for this (header offset 104, flagged unconfirmed across
several sections) directly: computed user_table_end (from the
already-solved channel+user table arithmetic, §6.2) for both USGS
files and compared to the literal value at header offset 104.
Magnetic_Data.gdb: user_table_end=579992, offset104's value
=580000 — off by exactly 8, not an exact match.
Radiometric_Data.gdb: user_table_end=933212, offset104
=933220 — off by exactly 8 again. A consistent 8-byte delta across
two files is suspicious enough to be something, but 580000 isn't
even page-aligned (580000/1024=566.406...), which is a bad sign for
"this is the blob region start" specifically (blobs are page-strided).
Went looking at other still-unlabeled header int32 fields instead of
stopping at the "close but not exact" offset 104 result. Header offset
108 in Magnetic_Data.gdb reads 567. 567*1024 (page_size) =
580608. Checked that exact byte offset: the CC CC 00 FF magic is
sitting right there. Did the same for Radiometric_Data.gdb: offset
108 reads 912, 912*1024=933888, and the magic is right there too.
Both real files, exact match, no slack. Header offset 108 is the
blob-region start, expressed as a page number; offset 104 was a red
herring (close by coincidence or shared rounding, not the answer).
Extended the same check to all 14 real GSQ files (looping
os.walk over samples/GSQ_Data, reading each file's first ~8MB,
computing header_offset_108 * header_offset_100 and checking for the
magic there) plus AG106386. All 15 matched, zero misses —
spanning every compression mode (NONE/SPEED/SIZE), chans_max
from 20 to 500, and three TEM system vendors across ~30 years. Combined
with the two USGS files, this field is now confirmed on 16 of 16
real files tested.
Then walked the entire chain, not just a sample, on 5 real
DB_COMP_NONE files — starting at the offset-108-derived position,
repeatedly reading a 48-byte header, checking the magic, and jumping
n_pages * page_size bytes to the next one, until either the magic
stopped matching (a real framing error) or the file ran out. Used file
seeks rather than loading multi-hundred-MB files into memory, after an
initial in-memory-buffer version of this check gave a false "the walk
just stopped" result on the two USGS files that turned out to be
nothing but the read buffer running out, not a real anomaly (caught
by re-running with a full streaming read instead of guessing the walk
had actually failed).
Result: all 5 files walk with zero framing errors, landing exactly on the file's true byte size:
| File | Size (bytes) | Blobs walked | Landed exactly on EOF? |
|---|---|---|---|
Magnetic_Data.gdb |
777,859,072 | 15,584 | yes |
Radiometric_Data.gdb |
706,635,776 | 23,079 | yes |
DB_EM_MountGordon_1003.gdb |
10,800,128 | 1,548 | yes |
DB_Mag_MountGordon_1003.gdb |
50,249,728 | 4,293 | yes |
DB_Mag_Elaine_1003.gdb |
2,249,728 | 777 | yes |
Two of these (the Mount Gordon files) are 1991 GSQ/Questem files,
independent of the USGS files in every way that matters (agency,
decade, vendor) and notable because their +16-onward trailer fields
(timestamp/scale/row-count/type) do not decode sensibly using the
same byte offsets that work perfectly for the 2020 USGS files (e.g. one
real header's +16 field read 0x80000000, not a plausible
timestamp) — yet the chain-walk itself (fields +0 through +12
only) still worked perfectly. Recorded honestly: the core fields are
far more stable across format vintages than the metadata trailer,
and the trailer's exact layout for older files is left as
[UNKNOWN] rather than forced to match the modern layout.
A full-file channel census (recording every distinct
blob_index % chans_max seen while walking Magnetic_Data.gdb
completely) found real blob runs for all 33 real channels already
identified in §6.2/NOTES.md, each with 630-641 blobs (matching the
~631 real lines from §6.3) except the known leftover/abandoned
channels (crap, deg, ch_11, __X, __Y, year_jd, crap2,
etc.), which had only 8-11 each — exactly the pattern expected if the
formula and the walk are both correct and nothing is being silently
skipped or double-counted.
One genuine loose end surfaced by the full census, not chased to a
conclusion: every channel's blob count included a handful of extra
blobs whose blob_index // chans_max decodes to an implausibly large
"line number" (up to ~1006, well past the real survey's line count),
and whose trailing fields show a repeating value (4670802) instead of
a valid GS_* code. Logged as [UNKNOWN] — plausibly some kind of
reserved/administrative "current value" blob outside the normal line
range (there's a suggestive but unconfirmed connection to
DB_CATEGORY_LINE_GROUP=200, already seen as a real line-record
category value during the Session 2 Mount Gordon investigation) — but
not pursued further given time, and not force-fit into the main
finding.
Compressed files — checked, only partially confirmed, left honest.
Ran the same chain-walk on AG106386 (DB_COMP_SIZE) and
DB_EM_293.gdb (DB_COMP_SPEED). AG106386 walked cleanly for 30
blobs with blob_index values 0..25 matching the known real channel
order (§6.5/NOTES.md) before moving to later lines — consistent with
the formula — and each blob's data was found to start with the
already-known 16-byte page-primitive magic (0f0efffe12345678, §6.5b)
immediately after the 48-byte blob header, i.e. the blob header wraps
the previously-solved compressed-page scheme rather than replacing it.
DB_EM_293.gdb (LZRW1) broke after only 3 blobs — a real framing
mismatch (+4 and +8 disagreed, 4 vs 1, whereas every
DB_COMP_NONE blob and the first 3 DB_COMP_SIZE blobs always had
these equal). Logged as a genuine, unresolved gap for DB_COMP_SPEED
specifically, not glossed over — likely a different page-accounting
convention for LZRW1-compressed blobs (compressed size not filling
whole pages the same way Size-mode's zlib pages do), but not
investigated further this session.
3.7 Reader and documentation updated¶
Extended reader/gdb_reader.py with BLOB_MAGIC/BLOB_HEADER_SIZE,
a BlobHeader dataclass, iter_blobs() (streaming chain walk from the
offset-108-derived start), find_blob(line_slot, channel_slot)
(computes the target index and walks until found or EOF), and
read_blob_values() (decodes a found blob's real data using the
owning channel's already-known GS_*/string-width type). Explicitly
raises rather than guesses for compressed files, per §3.6's honest
partial result. Wrote scripts/find_blob_index.py as a clean,
runnable, documented version of this session's discovery process (the
existing but incomplete scripts/locate_channel_data.py from the
killed session was left as-is rather than overwritten, since it's a
legitimate, if unfinished, independent piece of work). NOTES.md
updated with a new §6.6 and revisions to §6.4/§7/§8 reflecting all of
the above.
3.8 Coordinator hypothesis: is "stored raw vs. compressed" a per-channel row-count rule?¶
The coordinator raised a specific, testable hypothesis after reviewing
§6.5d's finding that two whole files declare DB_COMP_SPEED but
contain zero compressed data: is that actually a special case of a
more general per-channel rule, where a file's smaller-row-count
channels are always stored raw regardless of the file's declared
compression mode, even in files that genuinely do compress some of
their data? Pointed specifically at the Melinda Downs/Holroyd
River/Scrubby Knob files as a testbed, since §6.5d already established
they contain a real mix of compressed and stored-raw chunks.
First obstacle: chain-walking from §6.6 doesn't directly work for
Speed-mode blobs. Tried reusing iter_blobs(), but (as already
flagged as a known limitation in §6.6) the sequential n_pages-based
walk breaks after only a few blobs for Speed-mode files. Worked around
this by scanning for the two magics independently (CC CC 00 FF blob
headers, 0f0efffe12345678 LZRW1 chunk headers -- both already known,
reliable signatures from §6.6/§6.5b) and attributing each chunk to
whichever blob header immediately precedes it in the file (via
bisect over the sorted list of blob-header offsets) -- robust
regardless of the exact header size, which turned out to matter (see
below).
First real per-channel table, and the pattern jumped out
immediately. Ran this against DB_AGG_1213.gdb: tabulated, per
channel, how many of its real chunks are compressed vs. stored-raw,
plus the average/min/max decompressed_length. Two things were
obvious at a glance: (1) decompressed_length is identical across
every channel of the same dtype in the file, regardless of
compressed/stored-raw status (all GS_DOUBLE channels: 10344-14984
bytes; the one GS_LONG channel, flight: exactly half that, matching
4/8 byte type-width ratio) -- immediately suspicious for a
row-count-based hypothesis, since if row count varied enough to cause
a raw/compressed split, it should show up as varying chunk sizes
too. (2) the actual split is wildly uneven by channel identity:
RADAR is 0 compressed / 40 stored-raw; EASTING (same dtype, same
size range, same file) is 40 compressed / 0 stored-raw.
Why decompressed_length doesn't vary by channel: re-derived a fact
already on record. decompressed_length == row_count_for_that_line *
type_width, and row count is fixed per line, shared by every
channel on that line (the same column-alignment fact already
established in §6.4/§6.6 -- every channel's data for one line spans
the same fiducial range). So within one file, comparing
decompressed_length across channels is really just comparing type
widths, not row counts -- there is no meaningful per-channel
"how many rows does this channel have" distinct from "how many rows
does this LINE have," except via how many lines a channel happens to
be recorded on at all (already characterized in a different context,
§6.6's channel census). This alone is close to a direct refutation:
if two channels have the provably identical size distribution in the
same file, a pure size/row-count rule cannot put them on opposite
sides of a raw/compressed split -- and EASTING/RADAR do exactly
that.
Confirmed with real decoded values, not just the header
statistics. Decoded one real EASTING chunk (compressed,
chunk_length-12 was 70% of decompressed_length) and one real
RADAR chunk (stored-raw, chunk_length-12 == decompressed_length
exactly) from the same file. EASTING's values are a smooth,
tightly-clustered run (429762.19, 429761.06, 429759.93, ...,
differences of about 1.1 between consecutive points) -- exactly the
kind of data where consecutive float64 values share long common byte
prefixes, which is what LZRW1's short-window back-reference matching
can actually exploit. RADAR's values are also visually smooth at a
glance (88.08, 87.56, 86.47, 86.11, 85.96, ...) but evidently carry
enough real low-order sensor noise that LZRW1 found nothing worth
compressing -- the encoder tried, got no improvement, and (matching
Ross Williams' own reference FLAG_COPY semantics, already identified
in §6.5c/§6.5d) stored it raw instead rather than wasting space on
compressed-format overhead for no benefit.
Checked for a fixed per-channel policy too (a weaker, still
size-independent alternative) and ruled that out as well. Several
channels (DRAPESURFACE_FOURIER, and the _FOURIER_2p67/_EQUIV_2p67
gravity-correction channels) show a genuine mix of both outcomes
for the same channel across different lines (e.g.
GDD_FOURIER_2p67: 3 compressed, 37 stored-raw) -- direct evidence
this is a real per-block, data-dependent outcome at write time, not a
fixed rule keyed by channel identity either. Cross-checked against
DB_Mag_1213.gdb and DB_Mag_1212.gdb (same delivery, different
survey blocks): BAROMETER behaves similarly poorly in both, but the
four magnetic channels (RAWMAG/COMPMAG/DCMAG/LEVMAG) are
genuinely mixed in one file and always compressed in the other --
even "which channels compress well" isn't fixed across nearby files
from the same delivery, consistent with it being a property of the
actual recorded data, not the channel definition.
Checked the four original GEOTEM/Questem Speed files too, for
contrast. DB_EM_293.gdb, DB_Mag_293.gdb, DB_Mag_833.gdb: every
single channel in all three files is 100% compressed, zero stored-raw
chunks anywhere. This rules out yet another naive alternative ("Speed
mode always has some stored-raw chunks") -- whether stored-raw
chunks appear at all depends on whether any of that particular
delivery's actual data resists LZRW1, which apparently isn't true of
these older GEOTEM channels (per-gate EM decay curves, mostly smooth
monotonic-ish sequences within a gate) but is true of some of the
Melinda Downs AGG survey's derived correction channels.
A genuine by-product finding, not chased further this session:
compressed (DB_COMP_SPEED) blob headers appear to be 56 bytes,
not the 48 confirmed for DB_COMP_NONE in §6.6 -- the LZRW1 chunk's
own 16-byte magic consistently sits 56 bytes after the owning blob's
CC CC 00 FF magic in every compressed record checked, not 48.
Partially decoded by hand on one single-chunk blob (+24:
decompressed_length duplicated; +28: chunk_length+16; +40:
float64 1.0; +48: real row count; +52: GS_* type code) --
logged as a concrete lead for finishing §6.6's still-partial
DB_COMP_SPEED chain-walk story, not pursued further since it wasn't
needed to answer the compressibility question actually asked.
Result reported honestly: hypothesis tested and refuted, real cause
identified and demonstrated. The row-count-cutoff hypothesis does
not hold -- refuted by a direct structural argument (same-size chunks
of different channels land on both sides of the split in the same
file), not just an absence of correlation. The actual driver is
per-chunk data compressibility, a genuine content-dependent decision
made by the encoder at write time, consistent with (and now directly
demonstrating, not just inferring) Ross Williams' reference LZRW1
FLAG_COMPRESS/FLAG_COPY design. The original, separate §6.5d
observation (two whole files with zero compressed data, correlated
only with file size) is a different question this doesn't resolve --
left open, cross-referenced rather than conflated. NOTES.md updated
with a new §6.5e.
3.9 A large new real-file batch: Ontario Geological Survey, a .gdb/.geoh5 pair, and more GSQ¶
The coordinator relayed a batch of new real files the operator had
manually downloaded (same reasoning as the earlier GSQ zips -- a human
browser session gets past portal friction that automated fetches
can't; the files themselves are still just publicly downloaded data),
sitting in C:\Users\Joseph\Downloads\. Copied/extracted the
manageable ones into the repo's samples/ tree (left the multi-GB
SAMAGEM set in place, unpursued -- see below):
MLGRAV.gdb(126,034,944 bytes) andMLMAG.gdb(745,669,632 bytes) ->samples/ontario_GDS1251/-- Ontario Geological Survey GDS1251, Mozhabong Lake: magnetic + gravimetric (not gradiometer) data, a data type not previously in this project's sample set, and the first sample from a third independent agency (after USGS and GSQ).MLMAG.XYZ.txt(596,577,827 bytes) -- paired ASCII ground truth forMLMAG.gdbspecifically. Left inDownloads/and stream-read (Python file object, plainreadline()on the first several lines) rather than copied in full, the same technique already used for the ~1GB Georgetown.datfile in Session 2.collection (1).zip-> extracted justrm001141/DB_Rad_1141.gdb(9,657,344 bytes) andrm001141/DB_Mag_1141.gdb(40,083,456 bytes) via Python'szipfile(without extracting the full 59MB archive) tosamples/GSQ_Data/extracted/rm001141/-- another small GSQ survey (Fisher Creek) for cross-validation diversity.collection.zip(1.4GB) -> extracted justcr148832/East_Isa_VTEM_Inversion.gdb(342,972,416 bytes) andcr148832/East_Isa_VTEM_Inversion.geoh5(163,081,125 bytes), again viazipfilewithout extracting the full archive, tosamples/geoh5_east_isa/-- a genuine.gdb+.geoh5pair for the same real delivery (a modern airborne VTEM electromagnetic inversion, East Isa/Mount Isa region).- Deliberately not pursued:
SAMAGEM.zip(6.3GB uncompressed.gdb, GDS1089 Saganash Lake),SAMAGEM_CDI.zip(1.9GB derived CDI.gdb), and 5×SAMAGEM_L*.zip(paired ASCII ground truth, ~3.6GB each uncompressed) -- flagged by the coordinator as available but lower priority, and left that way given how much the rest of this batch already delivered; noted inNOTES.mdas a concrete, ready-to-use lead for anyone continuing this work.
3.10 Small GSQ pair (rm001141) -- routine cross-check, one real new data point¶
Ran the existing reader against DB_Rad_1141.gdb/DB_Mag_1141.gdb:
both parse cleanly (29 and 18 real channels respectively,
chans_max=50, comp_level=1). Ran the §6.5e-style
compressed/stored-raw chunk census against both and found zero
LZRW1 chunk-magic hits in either file, despite both declaring
DB_COMP_SPEED. Checked whether the data is present some other way:
find_blob()/read_blob_values() (with comp_level=0, i.e. treating
the blobs as plain) on Fid/Line for both files returns real, sane
values (Fid incrementing by 10 per row, Line a constant real line
number) -- confirmed both files store all their channel data
completely uncompressed, exactly like the two much larger
DB_Rad_1027.gdb/DB_Mag_1027.gdb files found in Session 2 §6.5d/2.11,
except these two are only 9.6MB and 40MB -- squarely inside the
2-12MB range of files that DO compress, not above it. This directly
refutes the file-size correlation §6.5d had explicitly flagged as
unconfirmed. Logged as NOTES.md §6.5f.
(One incidental fix needed along the way: find_blob(path, line_slot=0,
...) returned None for both files at first -- turned out line
numbering in this delivery starts at line slot 1, not 0, so there's no
real line-0 data to find. Not a bug in the reader; just picked the
wrong line_slot on the first attempt. Found the real starting line
slot by dumping the first several blob headers by hand and reading
their blob_index // chans_max.)
3.11 Ontario files break iter_blobs() immediately -- found and fixed a real one-line bug, with a much bigger payoff than expected¶
Ran the same header/symbol-table checks against MLGRAV.gdb and
MLMAG.gdb: both comp_level=0, chans_max=200. MLGRAV.gdb has a
genuinely new data type -- channels like grav_raw, FA_corr,
FA_anom, boug_corr267, boug_anom267 (free-air and Bouguer gravity
anomaly processing, not previously seen -- the Melinda Downs "AGG"
files from Session 2 are gravity gradiometer data, a different
measurement).
Ran the whole-file blob-chain walk (§6.6) against both files as a
routine sanity check before doing anything else with them.
MLGRAV.gdb got most of the way (11,870 of an eventual 11,891 blobs)
before hitting a real anomaly right near the end. MLMAG.gdb failed
immediately -- the very first blob the walker reached had
n_pages=2 (relative +4) but n_pages_dup=1 (relative +8), and
iter_blobs() (as written after §6.6) required these to match before
trusting a blob.
Rather than shrugging this off as "Ontario files are different,"
dumped the actual header bytes and checked which field, trusted alone,
actually lands on the next real blob header. n_pages (the first
field, not the duplicate) does -- exactly 2 pages later, and the next
real CC CC 00 FF magic is sitting right there. n_pages_dup is
simply wrong on this specific record, not a different-but-valid
convention. Checked what kind of record this actually is: blob_index
= 400004, and 400004 // 200 = 2000 -- the exact same class of
reserved/administrative blob (out-of-range line index, gs_type_code
reading the same 4670802 constant already flagged [UNKNOWN] in
§6.6, timestamp reading the same 0x80000000 sentinel) already
noted as a loose end in §6.6's original write-up. These administrative
records evidently just don't obey the "two fields agree" invariant
that happens to hold for every ordinary data blob checked so far --
the original iter_blobs() was over-fitted to that coincidence.
The fix: trust n_pages alone; drop the n_pages == n_pages_dup
requirement entirely. Re-ran the exact same whole-file walk against
both Ontario files: both now land exactly on their true file size
(11,891 blobs / 126,034,944 bytes for MLGRAV.gdb; 12,289 blobs /
745,669,632 bytes for MLMAG.gdb), with the previously-seen anomalies
(1 and 3 respectively) simply being administrative blobs that are now
correctly skipped over rather than causing a false stop.
Then, on a hunch, re-ran the identical fixed walker against every
other real file already in the project's sample set -- not just the
two that had just broken. This is where the payoff turned out to be
much bigger than "fixed two files": DB_EM_293.gdb (a real
DB_COMP_SPEED/LZRW1 file) went from walking cleanly for only 3 blobs
before breaking (the "genuine unresolved case" NOTES.md §6.6 had
explicitly flagged) to walking all 1,838 of its blobs with zero
errors, landing exactly on its true 12,370,944-byte size. Checked
every other real file in the same way -- 19 of 19 real files across
all three DB_COMP_* modes (NONE, SPEED, SIZE) now walk to an
exact byte-perfect EOF match. The earlier "compressed files only
partially verified" caveat throughout §6.6 was never a real limitation
of the file format -- it was an artifact of this project's own
over-strict sanity check, which happened to survive undetected on the
first DB_COMP_SPEED file tried (only 3 blobs in) but would always
have failed on a real administrative blob eventually. Fixed
reader/gdb_reader.py's iter_blobs() accordingly and updated
NOTES.md with a new §6.6b covering the complete, all-modes picture
(including the full 19-file table).
3.12 Wiring up compressed-blob data decoding (not just locating it) -- found a genuine third on-disk blob variant¶
With locating fully solved for all three modes, the natural next step
was decoding what's actually inside compressed blobs -- NOTES.md
still described read_blob_values() as raising NotImplementedError
for anything but DB_COMP_NONE. Worked this out directly against
AG106386 (DB_COMP_SIZE, zlib), since §6.5 already had known,
independently-verified ground truth to check against
(GA_project_number=5027, Fiducial=113120, 113140, ...).
- Found the compressed blob header is 56 bytes, not 48. Located
blob_index=0's header via the now-fixediter_blobs()/find_blob(), then searched nearby bytes for the already-known 16-byte page-primitive chunk magic (0f0efffe12345678, §6.5b/§6.5) -- found it at the header's relative+56, not+48. Decoding the zlib stream starting 16 bytes after that (i.e. atblob.offset + 72) reproduces the constant5027exactly, matching §6.5's real, independently-established ground truth. (First attempt tried+48, as the plainDB_COMP_NONEheader size, and failed with a zlib "incorrect header check" -- the wrong-offset failure mode itself was the clue that pointed at checking+56instead.) - Confirmed the same 56-byte header shape applies to real
DB_COMP_SPEED(LZRW1) blobs too, reusing the already-fully- validatedreader/lzrw1.pydecoder for the payload once the correct starting offset was known. - While testing this against
DB_EM_293.gdb(a real Speed-mode file), found a real, previously-unnoticed third blob variant. Some individual blobs inside this file have no chunk magic at all at the 56-byte-header position -- checked directly, not assumed, by searching a 200-byte window around one such blob for the 16-byte magic and finding nothing. Decoding this blob instead as if it were a plainDB_COMP_NONEblob (48-byte header, raw data immediately after) worked perfectly: real, sane, decreasing AMG-projected easting values (696511.0, 696501.0, -1e+32, 696491.0, ...) with the realrDUMMY=-1.0E32sentinel (vendor constant, §2) appearing exactly where a missing/no-fix sample would be expected. So a blob living inside an otherwise-genuinely-compressing file can be stored with no compression apparatus whatsoever -- a third real on-disk representation, alongside "genuinely compressed" and "chunk-wrapped but marked stored-raw" (§6.5e's finding). Updatedread_blob_values()to auto-detect this per blob (probe for the 16-byte magic; fall back to the plain layout if it's not there) rather than trusting the file's declaredcomp_levelat face value -- consistent with this project's running theme (first surfaced by the.grdCOMP_TYPE-lies-about-LZRW1 finding, then Seequent's own zlib-for-both-tiers documentation being wrong for Speed mode, and now a third instance of "a container-level claim doesn't guarantee anything about a specific piece of data inside it"). - Left honestly unresolved: multi-page compressed blobs
(
n_pages > 1). §6.5 already showed each 32768-byte page-sized slot independently starts its own zlib stream, but how successive pages within one logical blob are meant to concatenate wasn't worked out this session --read_blob_values()raisesNotImplementedErrorfor these rather than guessing.
Updated reader/gdb_reader.py: added COMPRESSED_BLOB_HEADER_SIZE,
imported zlib and the existing lzrw1 module, refactored the
value-decoding logic into a shared _decode_numeric_or_string() helper
used by both the plain and compressed paths, and extended
read_blob_values() with comp_level/page_size parameters and the
three-variant auto-detection described above. Re-verified all
previously-working cases (AG106386 zlib, Ontario/USGS plain data) still
decode correctly after the refactor.
3.13 Full value verification on a third agency, and the .geoh5 cross-check¶
Ontario (GDS1251), full value verification. With iter_blobs()
fixed (§3.11), ran find_blob(line_slot=0, ...) against MLMAG.gdb
for fiducial, x_nad83, y_nad83, mag_raw, and the string channel
line_number -- all five decoded exactly matching the real first three
rows of MLMAG.XYZ.txt (fiducial=68382.0, 68382.1, 68382.2;
x_nad83=389639.23, 389639.27, 389639.29; mag_raw=55364.83, 55364.58,
55364.32; line_number='1001' repeated), extending the complete
value-level ground-truth confirmation already done for USGS (Session 1)
and GSQ (Session 2, via ASEG-GDF2) to a third, unrelated agency.
MLGRAV.gdb's grav_raw/FA_anom/boug_anom267 decoded to
physically sane absolute-gravity/anomaly values (~981,700-982,700 mGal
raw; small residual anomalies) -- no independent ASCII ground truth was
available for this specific file, so this is a plausibility check
rather than a digit-for-digit match, logged as such.
.geoh5 cross-check. Before touching the paired
East_Isa_VTEM_Inversion.geoh5, checked geoh5py's license directly
rather than trusting the coordinator's characterization at face value
-- PyPI's JSON metadata API returned an empty license field, so
fetched pyproject.toml straight from github.com/MiraGeoscience/
geoh5py instead: license = "LGPL-3.0-or-later". This is a
completely separate, independently-published, open-source library for
a different, openly-documented Seequent format (not the
geosoft/gxapi/gxpy proprietary package this project is barred
from touching), so pip install geoh5py doesn't touch the hard
constraint. Installed it and opened the file with
geoh5py.workspace.Workspace.
Found 258 DrapeModel objects (2D inversion cross-sections), each
named after a real survey line (L1000, L1010, ..., L4021) -- no
raw point-data object with individual channel values in this
particular .geoh5 (only inversion results: the drape models, a DEM
surface, 2 geology images), so a true numeric cross-check wasn't
possible with this file specifically. Extracted the sorted set of all
258 DrapeModel names and independently scanned the paired .gdb
file's own line symbol table (same 128-byte-stride,
category-100-filtered technique as §6.3/§6.6) for real line-name-
shaped tokens: also exactly 258, and the two sets are identical
-- zero names in either set missing from the other. Two
structurally unrelated container formats (this project's own
from-scratch .gdb parser, and Mira Geoscience's independent, open,
HDF5-based .geoh5 reader) agree exactly on the real content of the
same delivery.
Also ran find_blob()/read_blob_values() against
East_Isa_VTEM_Inversion.gdb's own channels (it's DB_COMP_NONE, and
walks perfectly per §6.6b's table): UTMX/UTMY decode to smoothly-
varying real UTM coordinates, RESDATA (raw apparent resistivity feeding
the inversion) decodes to values in the 0.03-0.07 ohm-m range --
consistent with the historically conductive Mount Isa mineralization
this survey covers -- and the array channel DEP_BOT (depth to bottom
of each of 24 inverted layers) decodes to a real, monotonically
increasing depth profile (4.0, 8.495, 13.547, 19.225, 25.606, 32.777,
...) with row_count (42,792) exactly equal to 1,783 stations × 24
layers -- reconfirming the VA/array-channel row-count semantics first
established in §6.2b, now on a fourth real array-channel example (this
file has 3: CHA_CROPP, DEP_BOT, RHO_CROPP, all array_width=24).
This file also has the rare f0f0f0f0 header-signature variant first
seen exactly once in Session 2 (DB_Mag_Elaine_1003.gdb) -- a second,
completely unrelated real occurrence of the identical 4 bytes, upgrading
it from "one-off" to "real, recurring, still-unexplained variant."
Updated NOTES.md with new §6.6b and §6.6c sections, revised the §6.1
f0f0f0f0 note, added a new §5 sample-file table block, added S14/S15
to the sources table, and rewrote the "Natural next steps" list to
reflect everything above.
3.14 Coordinator self-assessment ask, and the answer: multi-page compressed blob decoding¶
The coordinator asked for an honest, prioritized read on what's left
in NOTES.md's own open items now that the big structural questions
are solved, explicitly ruling writing/mutation support out of scope
going forward. Went through every listed open item (unknown header
fields, line-table layout, REG/coordinate-system parsing, the
unexplained oddities, the newest-round threads, and whether SAMAGEM is
worth returning to) and reported back: everything is now either
decorative/non-blocking (the unknown header fields, the f0f0f0f0
variant, the UTF-16 path, the administrative blobs -- all already
safely handled or simply curiosities) or a bounded, optional follow-up
(REG parsing, the line-table layout) -- except multi-page
compressed blob decoding, which was still a genuine, concrete,
functional gap: read_blob_values() explicitly raised
NotImplementedError for any compressed blob with n_pages > 1,
which is a real, common case in real files (some right in this
project's own sample set have blobs up to 47 pages). The coordinator
agreed and asked to tackle that one next, with a specific steer: don't
assume the single-page case just generalizes to multi-page -- verify
it first, specifically -- and use SAMAGEM_CDI.gdb opportunistically
if it turns out to be a natural real-world testbed, without feeling
obligated to touch the rest of the multi-GB SAMAGEM set otherwise.
3.15 Testing the two competing hypotheses for multi-page blobs, directly¶
Two structurally different possibilities were worth distinguishing before writing any code: (a) each page-sized slot within a multi-page blob independently starts a fresh chunk (its own 16-byte magic + a new compressed stream, chained page-to-page like a miniature version of the outer blob chain), or (b) the whole blob is one compressed stream that simply spans across page boundaries with no re-framing at all. §6.5's original observation (every page-sized slot in the general data region independently starts its own zlib stream) was consistent with either, since that scan never specifically checked within one already-identified multi-page blob.
Took a real, already-located 2-page DB_COMP_SIZE blob (channel 9 =
Easting, line 0, AG106386) and checked directly: does the second
page (blob.offset + page_size) start with the 16-byte page magic the
way the first page does? No -- the first bytes of page 1 are 4b
bc 30 c6 a6 4f 84 06 ..., nothing like 0f 0e ff fe 12 34 56 78.
That single check rules out hypothesis (a). Tested (b) directly by
reading the entire n_pages*page_size span (minus the 72-byte
header+magic prefix on page 0) as one byte string and handing all of
it to zlib.decompressobj().decompress(...) in a single call: decoded
cleanly to 84,432 bytes (10,554 real float64 values, a smooth
easting-coordinate profile), with decompressobj correctly finding
the real end of the zlib stream partway through the buffer and
reporting the rest as harmless trailing page padding (unused_data).
Hypothesis (b) confirmed directly, not by elimination alone.
Checked it holds at real scale, not just on one small example.
Found a real 36-page blob (LEI_Depth, an array_width=30 channel,
same file) and applied the identical technique: decompressed cleanly
to exactly 10,554 stations x 30 layers x 4 bytes = 1,266,480 bytes.
A real detour, caught and corrected rather than reported as a
finding: the first attempt at this 36-page test used the
neighboring channel, LEI_Conductivity, and assumed (without
checking) that it shared LEI_Depth's GS_DOUBLE declared type.
Decoding its bytes as float64 produced obviously-wrong,
denormalized-looking tiny numbers (~1e-17 to ~1e-31) -- briefly
looked like a real, interesting "compressed array channels use a
narrower on-disk width than declared" finding, until re-checking the
channel's own actual dtype_code directly: it's 4 (GS_FLOAT), not
5 (GS_DOUBLE) -- a different channel with a different real type,
not a hidden compression-specific width convention. Decoding it as
GS_FLOAT (which the existing, unmodified dtype-driven decode logic
already does automatically) gives clean, physically sane, smoothly-
varying conductivity values with no garbage. Recorded honestly in
NOTES.md as a self-caught false alarm from careless manual probing,
not a reader bug -- the actual library code was never wrong here, only
this session's ad hoc investigation script briefly was.
Re-ran LEI_Depth (correctly, GS_DOUBLE) properly: values matched
exactly the known real depth profile from an earlier session's
independent ASEG-GDF2 ground truth (0.0, 3.0, 6.3, 9.9, 13.9, 18.3,
23.2, 28.5, ...), repeated identically once per station -- physically
correct, since depth-to-layer is the same fixed profile for every
sounding in a layered-earth inversion, only conductivity varies row to
row.
Confirmed for DB_COMP_SPEED (LZRW1) too, on a real 2-page blob
(Northing_AMGz55, DB_EM_293.gdb): read the full 2-page span past
the 56-byte header and handed it to the unmodified
lzrw1.parse_chunk_header()/decode_speed_chunk() -- these already
worked from the chunk's own recorded decompressed_length/
chunk_length fields rather than any page-count assumption, so
nothing about them needed to change at all. Decoded to 1,120 real,
sane northing values with real rDUMMY sentinels in the right places.
3.16 The fix, and a stress test¶
The actual code change was smaller than the investigation that led to
it: removed the if blob.n_pages != 1: raise NotImplementedError
guard in reader/gdb_reader.py's read_blob_values(). The
surrounding code already computed blob.n_pages * page_size -
COMPRESSED_BLOB_HEADER_SIZE and read that whole span -- that was
already the correct multi-page behavior, just gated off behind an
overly cautious check written before it had been verified.
Stress-tested across 4 real compressed files (AG106386,
DB_EM_293.gdb, DB_EM_833.gdb, DB_AGG_1213.gdb), 400 blobs each
(1,599 total, n_pages ranging 1-47): every blob with a sane header
decoded without error. The only failures were the already-known
reserved/administrative blobs (row_count < 0, flagged [UNKNOWN]
since §6.6), which crashed with a confusing low-level struct.error
("read length must be non-negative or -1") rather than a clear
message -- fixed alongside the main change by raising a clean
ValueError naming them explicitly as non-data administrative blobs
in both places read_blob_values() does a plain byte read (the
comp_level==0 path and the "bare"-blob fallback path). Re-ran the
stress test: zero unexpected errors.
Also updated the CLI demo (gdb_reader.py --main--) which still
printed a stale "multi-page compressed blobs are not yet handled"
message even after the underlying function was fixed -- corrected to
reflect the real, current capability.
3.17 Opportunistic large-file check: SAMAGEM_CDI.gdb¶
Per the coordinator's specific allowance, extracted the
already-downloaded SAMAGEM_CDI.zip (1.2GB) to get at
SAMAGEM_CDI.gdb (1,930,303,488 bytes, GDS1089 Saganash Lake derived
conductivity-depth-imaging database) and checked its header: comp_
level=0 (DB_COMP_NONE) -- so it didn't end up testing the
compressed multi-page fix specifically after all. Rather than
discarding it, ran the whole-file blob-chain walk against it anyway
as a scale check, since it was already in hand: 6,066 blobs, landing
exactly on the true 1,930,303,488-byte file size -- by far the
largest file this project has walked end-to-end (previous largest was
Magnetic_Data.gdb at 778MB), a useful confirmation the model holds
at nearly 2GB. Also noted a real array_width=50 channel
(em_z_final_off) -- the largest VA/array width seen in this project
so far (previously 24 and 30). Did not touch SAMAGEM.zip or the
SAMAGEM_L*.zip ground-truth files, per the coordinator's explicit
"don't feel obligated" framing for the rest of that set.
3.18 Write-up¶
Updated NOTES.md with a new §6.6d covering the full multi-page
investigation and result, corrected the now-stale "multi-page
compressed blobs are a remaining gap" language in §6.6b's bullet list,
the §7 working-reader summary, and the "Natural next steps" list
(item 1 now fully resolved rather than partially), added
SAMAGEM_CDI.gdb to the §5 sample-file table, and updated the
sample-file-count language throughout (23 real files now validated,
up to 1.93GB). Bottom line: "read any channel, any file, any
compression mode" -- the coordinator's own framing of the priority --
is now genuinely complete for single- and multi-page blobs alike,
across all three DB_COMP_* modes.
3.19 A standalone SPEC.md, and a real counting error caught while writing it¶
The coordinator asked for a consolidation/writing task, not more
research: a standalone reference specification for the .gdb format,
separate from NOTES.md/LOG.md (both chronological research logs,
not references), organized by the file's actual structural pieces
(header, symbol table, blob index, compression schemes) rather than by
the order any of it was discovered in, reusing the same confidence
markers throughout for consistency, and citing NOTES.md section
numbers rather than re-deriving evidence inline.
Wrote SPEC.md at the repo root: conceptual model; file header (byte
offset table); symbol table (shared structure, then channel/line/user
record layouts separately); data types/formats/dummy values; VA/array
channels; the blob index (addressing formula, chain-walking, the plain
and compressed blob header layouts, the "bare blob" and multi-page
findings); the three compression schemes and their exact on-disk
framing; the .geoh5 cross-validation; a consolidated list of known
real-world oddities that are safely handled but not understood; a
reference-implementation pointer; and an explicit "not covered" section
(REG/coordinate-system parsing, full line-table layout, and — per the
coordinator's separate note this session — writing/mutation, now
explicitly out of scope for the whole project going forward).
Caught and fixed a real arithmetic error while cross-checking the
draft against NOTES.md before finalizing, rather than propagating
it: NOTES.md's own "working reader" section claimed validation
against "23 independent real files (2 USGS + 17 GSQ + 4 Ontario)" --
but recounting the actual Ontario sample set (MLGRAV.gdb,
MLMAG.gdb, SAMAGEM_CDI.gdb) gives 3 files, not 4, making the real
total 22, not 23. This was a genuine bookkeeping slip from an earlier
session-3 edit (probably drafted before SAMAGEM_CDI.gdb was added
and never re-summed), not a new finding -- fixed the count in both
NOTES.md and the SPEC.md draft rather than let a wrong number carry
into a document explicitly meant to be an authoritative reference.
Also refreshed README.md, which had gone stale (still describing the
.gdb reader as "partial" and not decoding channel data at all, and
the sample-file list as 10 files from 2 agencies) -- pointed it at
SPEC.md as the new primary entry point, kept NOTES.md described as
the supporting evidence trail, and brought the reader-capability and
sample-file descriptions up to date with everything since Session 2.
3.20 REG/coordinate-system (IPJ/map-projection) metadata -- found and partially decoded¶
The coordinator flagged this as "the one gap from the spec review
that's actually tractable right now" (as opposed to a few others that
need a new sample file with a specific feature not currently in hand)
and asked for exploratory search rather than assuming a structure up
front: real airborne surveys are georeferenced, this project already
has real files whose sidecar metadata literally names the projection
used (the Georgetown ASEG-GDF2 .prj file), and a projection-name
dictionary plus IPJ/__dbreg-style registry markers had already been
incidentally spotted once, unexplained, in Session 2 (LOG.md §2.3).
First pass: re-examined the Session-2 find directly. Dumped the
known blob region (DB_Mag_833.gdb, roughly offset 316000-343000) and
extracted every printable string. Confirmed it's exactly what Session 2
guessed: a huge, generic, Geosoft-bundled catalog of named projections
(hundreds of US SPCS83 state-plane zones, Swedish/Taiwanese/historical
Texas zones, etc.) -- clearly shipped identically with every database,
not survey-specific. Also found a much more interesting real,
survey-specific registry entry in the same region:
"?|IPJ_Easting_AGD66:Northing_AGD66" -- a literal key naming the
real channel pair this database's working projection applies to
(Easting_AGD66/Northing_AGD66 are this file's actual real channel
names). Dumped the bytes right after this key: mostly zero padding,
then a short int32 (0x03ac = 940, unclear meaning) followed by a long
run of what are almost certainly raw 64-bit in-memory pointers
(structured 8-byte values sharing a 0x1211/0x1212-prefixed high
half, consistent with a live Windows process's heap addresses) rather
than portable data -- a real, if unglamorous, finding: this specific
registry region looks like it's carrying serialized application
state (nearby strings: "Display List", "Database Extension
Objects"), not a clean, portable coordinate-system record. Recorded
honestly rather than forced into a projection-parameter narrative that
didn't fit the actual bytes.
The real breakthrough came from a different angle: numeric ground
truth, not string search. The Georgetown .prj sidecar
(samples/GSQ_Data/extracted/georgetown_ascii/Northern Georgetown.prj,
already in hand since Session 2) is a single-line ASEG-GDF2 PROJ
record: PROJGDA2020 / MGA zone 54 GDA2020 6378137
0.0818191910428158 0 Transverse Mercator 0 141 0.9996 500000
10000000 -- i.e. explicit, exact values for ellipsoid semi-major axis,
eccentricity, central meridian, scale factor, false easting, and false
northing. Rather than guessing at a binary layout, searched
AG106386_Northern Georgetown_Conductivity.gdb (the paired .gdb for
this exact survey) for these six specific float64 values, literally,
as raw bytes -- and found every one of them, repeated across the
file at regular ~32768-byte (page_size) intervals, with central
meridian/scale/false easting/false northing sitting at
consistent small deltas from each other (+24/+8/+8 bytes) --
clearly a real, structured record, not coincidental noise.
Traced one occurrence back to its owning blob and got the other
half of the picture. The nearest preceding CC CC 00 FF blob magic
sits 596 bytes before the hit. Its header decodes to blob_index=
500006 -- with chans_max=500, that's line_slot=1000, exactly the
same out-of-range-line-index "reserved/administrative blob" signature
already flagged [UNKNOWN] in NOTES.md §6.4/§9 (constant
gs_type_code, implausible line number). This directly explains a
previously-unexplained real anomaly: at least some of those
administrative blobs are per-database coordinate-system metadata,
reached through the exact same blob-chain mechanism as real data, just
living in a reserved line-index range used as a metadata namespace
rather than real survey lines.
Dumped and read the full blob by hand. Found real, human-readable
strings: "IPJ" (the literal object-type tag), a repeating 4-byte
marker " JPI" immediately followed by int32(1) and a NUL-terminated
name string (confirmed exactly preceding "WGS 84 / UTM zone 54S",
the working projection actually used in the raw survey data --
different in datum-naming convention but numerically identical to the
delivered .prj's "GDA2020 / MGA zone 54", a completely normal
real-world discrepancy between an internal working CRS and a delivered
client-specified one), plus further un-tagged strings for the datum
name ("WGS 84", appearing multiple times) nearby. Also found more
of the same in-memory-pointer-shaped byte patterns already flagged as
suspect in the DB_Mag_833.gdb investigation above -- same phenomenon,
now recognizable rather than a fresh surprise.
Generalized the search and confirmed on two more independent
agencies, both geographically correct. Built a small reusable
technique: walk iter_blobs(), flag any blob whose line_slot is far
outside the file's real line count, then probe the first ~4KB of each
flagged blob for the literal bytes IPJ. Ran this against
Magnetic_Data.gdb (USGS): found 1726 administrative-blob
candidates in total, but only 3 actually contain IPJ content --
confirming this is a real, if not exclusive, purpose for these blobs
(most of the 1726 instead start with a different tag, "REG ",
followed by what looks like real float64-shaped survey data rather
than coordinate-system metadata -- a separate, unexplored mechanism,
not chased further). The 3 real IPJ blobs found in Magnetic_Data.gdb
contain: "NAD83 / UTM zone 11N", "NAD83", "GRS 1980" (the correct
NAD83 reference ellipsoid), and "NAD83 to WGS 84 (1)" (a named datum
transformation) -- and UTM zone 11N is exactly the real-world-correct
zone for this survey's actual location (southeast Mojave Desert,
~115°W). Repeated the identical technique against MLMAG.gdb
(Ontario): found "NAD83 / UTM zone 17N" plus the same "GRS 1980"/
"NAD83 to WGS 84 (1)" companions -- UTM zone 17N being exactly
correct for Mozhabong Lake, northeastern Ontario. Three unrelated
agencies, three geographically correct results, using nothing but the
literal bytes already in each file.
Honest stopping point. Did not attempt to fully map the record
byte-by-byte (most of the ~700 bytes examined by hand around each hit
remain unidentified), did not determine how the projection/ellipsoid/
transformation sub-objects are delimited from each other within one
blob, and did not chase the "REG "-tagged majority of administrative
blobs found in Magnetic_Data.gdb at all. This matches exactly the
kind of partially-open outcome the coordinator said would be fine --
recorded as such in NOTES.md (§6.7, a new subsection) rather than
either overclaiming a full decode or silently dropping the thread.
Updated SPEC.md to add a proper §8 section for this (renumbering
the sections after it) and to remove the now-stale "never attempted"
framing that was in the header section.
3.21 Following up on the "REG " lead: Geosoft Desktop's own settings/processing-history registry¶
Direct follow-up per the coordinator: most out-of-range-line_slot
administrative blobs found while investigating IPJ (§3.20) are
not projection records -- they start with a different 4-byte tag,
"REG ", left explicitly unexplored. The coordinator suggested three
angles and left the order up to me: (1) check for the same "tag +
marker + name" micro-pattern found for IPJ; (2) check whether
vendor-published DB_* constant names (already in hand from
gxapi/__init__.py) show up as readable strings inside these blobs;
(3) cross-reference real per-survey metadata sidecars already in hand
(Readme.txt, .des files, etc.) the way the .prj sidecar's exact
numeric values cracked open the IPJ blob.
Angle 1 first, since it was the most direct. Dumped a real "REG "
blob byte-for-byte (Magnetic_Data.gdb, offset 115064832). Found the
same general tagged-object convention as IPJ: 48-byte blob header,
then a NUL-terminated short name ("REG\0"), then a space-padded
4-byte FourCC-style tag ("REG ") + int32(2) + int32(1) + a
recurring separator constant, then a nested tag ("VV " --
Geosoft's own "VV" vector-value object, named in vendor source),
itself sometimes followed by a further named sub-field ("CLASS",
"MAKE", both introduced the identical way). Confirmed this is a
real, reused framing convention -- not unique to IPJ -- but didn't
attempt to map every byte.
Angle 2 turned out to be mostly a miss, but pointed at something
better. No literal DB_SYMB_*/DB_CHAN_*-style vendor constant
names were found as readable text in any REG blob. Broadened the
search instead of stopping there: extracted every printable string of
5+ characters from all 466 real "REG "-tagged blobs found in
Magnetic_Data.gdb (walking iter_blobs(), flagging line_slot
values far outside the file's real line count -- the same technique
already built for the IPJ search) and counted them by frequency.
This is where the real payoff was: a completely different, much richer
vocabulary of registry key names showed up immediately --
UNITS, LABEL, CLASS, MAKER, FORMULA, and clearly
Oasis-montaj-tool-specific keys like LOOKUPDBCH.REFCH,
MATHEXPRESSIONBUILDER.CHANNELEXPRESSIONFILE -- plus real dates
(2019/12/18, 2019/12/20, 2019/12/29, 2020/01/22) and a literal
GX tool identifier string
(geogxnet.dll(Geosoft.GX.MathExpressionBuilder.MathExpressionBuilder;
RunChannel)). This alone was enough to recognize these blobs as some
kind of real settings/history log, not noise -- worth pinning down
further before writing anything up.
Angle 3 confirmed it decisively, in both directions. One of the
extracted strings was an actual user-entered formula:
MATHEXPRESSIONBUILDER.CHANNELINPUTBOX="ch_9=comp_mag - ch_8;
ch_9=ch_9 + 48066.0;". Went straight to this exact file's own
Readme.txt (already in hand since Session 1, re-read rather than
assumed) to check: it states plainly that "Magnetic data were
processed by EDCON-PRJ, Inc. and include corrections for diurnal
variations of the Earth's magnetic field, magnetic field of the
aircraft, tie-line leveled, micro-leveled..." -- the recovered formula
(subtract a base/diurnal channel, add a constant leveling offset) is
exactly the shape of correction that sentence describes. This is the
single clearest piece of evidence in this whole investigation that
these blobs are real, meaningful records and not junk. The same
Readme.txt also states "Data are in the World Geodetic System 1984
(WGS84) and also in UTM projection, Zone 11, North American Datum
1983 (NAD83)" -- which matches, a second and completely independent
way, both the binary IPJ blob from §3.20 and a newly-found textual
serialization of the identical projection settings sitting inside one
of these REG blobs (literal strings "NAD83 / UTM zone 11N",
NAD83,6378137,0.0818191910428158,0, plus internal key names --
_PJ_NAME, _PJ_ELLIPSOID, _PJ_DATUM_TRANSFORM, _PJ_PROJECTION,
_PJ_UNITS, _PJ_X, _PJ_Y, _PJ_IPJ -- that plausibly explain how
the binary IPJ record's own sub-fields are keyed internally, a real
bonus completion of §3.20's open "how are sub-objects delimited"
question, found via a completely different lead).
A few more specific fields dug into by hand, for completeness:
- LABEL fields, where populated, hold real provenance strings:
"Source: .\delete.gdb" and "Source: .\gps\mag_gps.gdb" --
matching real intermediate/working files also referenced by the
LOOKUPDBCH.DB=".\delete.gdb" GX-tool-parameter string found
separately -- a self-consistent processing-lineage trail.
- UNITS fields, where populated, use a different, more compact
encoding than the GX-tool-parameter style (KEY="value"\r\n):
UNITS\x00dega,1 and UNITS\x00m,1 -- i.e. plain unit codes
(decimal degrees, meters) for real channels, found by specifically
searching for UNITS occurrences with non-empty content following
them (many UNITS/CLASS slots turned out to be bare placeholders,
just the key name followed immediately by zero padding -- a real,
if unexplained, distinction between populated and empty registry
entries).
- MAKER did not turn up a readable text value in any instance
checked -- instead, right where a value might be expected, the bytes
decode as yet another "MAKE" FourCC-style tag, suggesting MAKER
introduces a nested sub-object rather than holding a flat string,
consistent with the same general tagged-object convention as
everything else here. Not chased further.
Honestly scoped, not overclaimed. This was all done against a
single file (Magnetic_Data.gdb, USGS) -- the technique itself
(flag out-of-range line_slot blobs, search their content for
readable text) is trivially reproducible on GSQ/Ontario files, but
doing so was left as a natural next step rather than squeezed into
this same round, since the USGS file alone already produced more than
enough to write up honestly, and confirming it generalizes to other
agencies wasn't necessary to answer the question actually asked (what
do the "REG " blobs contain). No byte-exact record parser was
written -- everything above came from targeted string search and hand
inspection of surrounding bytes, not a generalized decoder.
Updated NOTES.md with a new §6.8, revised the §6.7/§8/"Natural next
steps" text that had called this thread unexplored, and updated
SPEC.md with a new §9 (renumbering the sections after it again) plus
a small independent fix: found and corrected two pre-existing SPEC.md
cross-reference bugs unrelated to this round's renumbering (the
VA/array-channel field table in §3 pointed at §8 instead of §5, likely
a copy-paste slip from the original draft) while doing the consistency
pass this round's renumbering required anyway.
3.22 Checking whether the REG registry finding generalizes -- confirmed on both other agencies¶
The coordinator asked directly: does the REG registry finding
(§3.21) hold up on GSQ and Ontario files, or was it a USGS-specific
coincidence? Explicit instruction not to assume it holds -- actually
check, same technique, real files.
GSQ: ran the identical scan against AG106386_Northern
Georgetown_Conductivity.gdb (a second, independent GSQ file --
already used for IPJ/blob-index work in earlier sessions, but never
scanned for REG content specifically). Found the same tag
distribution ("REG " dominant among administrative blobs, IPJ
present as a minority) and the same kind of content: real GX tool
run records (newchan.gx/"New channel", copy.gx/"Copy channel",
the same MathExpressionBuilder tool as the USGS file), real
parameter strings for channel creation (NEWCHAN.DTYPE="Double",
NEWCHAN.DISPDIG="4", etc. -- a genuinely useful new lead for a few
of this project's still-unresolved channel-record display fields),
and two real formulas referencing this exact file's own real channels
(a YYYYMMDD date-formatting expression using a real Date_ channel;
a grid-math expression referencing a real external SRTM elevation
grid file).
The GSQ file also gave a stronger numeric confirmation than the
original USGS result. A textual IPJ serialization inside one of
its REG blobs reads "GDA2020 / UTM zone 54S",
GDA2020,6378137,0.0818191910428158,0, "Transverse Mercator",0,141,
0.9996,500000,10000000. Checked digit-by-digit against
Northern Georgetown.prj (the same ASEG-GDF2 sidecar already used to
crack open the binary IPJ blob in §3.20): every one of the six
numeric parameters matches exactly -- not just consistent order of
magnitude, a literal digit-for-digit match between an independently-
sourced plain-text industry-standard file and a completely separate
internal registry string inside the binary .gdb. Recorded as the
single strongest confirmation in this whole REG/IPJ investigation.
Ontario: ran the same scan against both MLMAG.gdb and
MLGRAV.gdb. Same tag distribution again. Same category of content:
UNITS/LABEL/CLASS/FORMULA keys, a textual IPJ serialization
("NAD83 / UTM zone 17N", NAD83,6378137,0.0818191910428158,0,
"Transverse Mercator",0,-81,0.9996,500000,0), and real formula
fragments (mag_igrf+mag_tlcor, floor(Line_number)) referencing
real channels -- checked directly against MLMAG.gdb's own channel
list (read_channels(), already validated in an earlier round):
mag_diurn, mag_igrf, mag_lev, mag_gsclevel, line_part are
all real, exact channel names in this file.
Two genuinely new things found on Ontario, not seen on USGS or
GSQ:
- The projection parameters correctly reflect the northern-
hemisphere UTM convention (false northing 0, not the 10000000
seen on the two southern-hemisphere Australian numbers) and the
central meridian (-81) is the real, independently, publicly
checkable geodetic value for UTM zone 17N -- an objective fact, not
something taken on the file's own word.
- Literal vendor-published constant names actually appear as
readable text: DB_CHAN_X and DB_CHAN_Y, sitting right next to
an IPJ_x_nad83:y_nad83 registry key, matching DB_CHAN_X=0
DB_CHAN_Y=1 from the vendor's own published source (already on
record, NOTES.md §2) exactly. This directly resolves "Angle 2"
from §3.21, which had come up empty on the one USGS file checked --
the miss was specific to that file, not a real absence from the
format in general.
Honestly scoped gap, not smoothed over: no dedicated GDS1251
processing-history sidecar (the Ontario equivalent of USGS's
Readme.txt or GSQ's .prj) was found or available this round --
checked the Downloads folder for anything resembling one and found
nothing clearly relevant (one large, ambiguously-named PDF turned up
in a search but had no clear connection to this specific survey
delivery and was not opened, out of appropriate caution about reading
arbitrary files from the user's general-purpose personal Downloads
folder without a real signal they're relevant to this task). The
cross-reference for Ontario is consequently a notch weaker than the
USGS/GSQ cases -- real channel-name matches against this project's
own already-validated read_channels() output, plus an independently
checkable public geodetic fact (UTM zone 17N's true central meridian),
rather than a literal third-party document quote. Recorded as exactly
that: real, but a different (and slightly weaker) kind of
confirmation than the other two agencies got, not glossed over as
equivalent.
Result: generalizes completely, no exceptions found. All three
agencies show the identical tagged-object framing and the same
category of real content. Updated NOTES.md §6.8 with a full "UPDATE"
section covering all of the above, updated the "Natural next steps"
entry that had flagged this as unconfirmed, and updated SPEC.md §9
to describe the finding as agency-general rather than USGS-specific.
3.23 Full-corpus sanity pass: everything, together, at once¶
The coordinator asked for something different from every previous
round this session: not a new targeted investigation, but a broad
sanity pass -- run the complete reader (header, symbol table,
blob-index/chain walk, data decoding across all compression modes,
VA/array channels, and the REG/IPJ registry scan) against every real
.gdb file collected so far, all three agencies, not a hand-picked
sample. The explicit concern: something that worked in isolation
during earlier rounds might break, or look subtly wrong, once combined
with a different file's quirks that hadn't been hit in the same run
before.
Method. Wrote scripts/full_corpus_sanity_check.py: discovers
every .gdb under samples/ via glob (not a fixed list -- so it
automatically covers anything added since), and for each one runs the
magic/header check, full channel/line symbol-table decode, a
whole-file blob-chain walk with an exact-EOF check, decodes the first
five real channels' actual data for the first real (non-
administrative) line found, and runs the REG/IPJ admin-blob scan.
Ran it against all 22 real files this project has -- the complete
corpus, 2MB to 1.93GB, all three DB_COMP_* modes, all three
agencies.
Headline result: zero exceptions, zero crashes, on all 22 files. Every blob-chain walk lands exactly on the true file size. Every sampled channel decodes to finite, physically sane values matching each survey's real known location (correct-sign, correct-magnitude coordinates for every region) and real known physical quantity (realistic raw magnetic totals, sane radiometric percentages, sane VTEM/EM decay and depth-of-investigation values). This alone answers the question actually asked: yes, everything genuinely holds together when run as one combined pass, not just pairwise.
A real near-miss in this round's own script, caught before it
became a false report. The first version capped the admin-blob scan
at 3,000 blobs per file for speed. Several real files have far more
total blobs than that (Radiometric_Data.gdb: 23,079) -- meaning the
cap could be exhausted by ordinary survey-line blobs before the walk
ever reaches the out-of-range-line_slot administrative region. This
produced a spurious REG=0, IPJ=0 result for Magnetic_Data.gdb --
the exact file the whole REG investigation (§3.21) was built on, which
really has 466 real REG blobs. Noticed the contradiction against
already-established results rather than trusting the new script's
output blindly, and re-ran the admin-blob scan with the cap removed
(scripts/reg_ipj_full_scan.py), which correctly recovers all 466.
Logged as a real methodological gap in this very round, not silently
patched.
New finding 1: the first non-GS_DOUBLE/GS_FLOAT array channel in
the whole project, hiding in a file examined since Session 1.
Re-decoding every channel of Radiometric_Data.gdb (rather than the
handful spot-checked in earlier rounds) surfaced ISPD and ISPU:
GS_USHORT, array_width=512. Decoded the real data directly:
row_count=145,408 is exactly 284 stations x 512 channels, and the
first station's 512-element vector is a textbook real airborne
gamma-ray energy spectrum -- zero counts below the detector's energy
threshold, a sharp rise to a peak around channel 10-14, then a smooth
physically realistic decay. LOG.md Session 1 §1.17 had already
logged ISPD/ISPU as ordinary scalar GS_USHORT channels -- correct
about the type, but the array-width field was simply never checked
for these two specific channels until this pass re-decoded everything
at once. Directly resolves an item that was still open on this
document's own "Natural next steps" list.
New finding 2: REG/IPJ administrative content is real but not
universal, and the pattern of its absence is suggestive. The
uncapped admin-blob scan found zero REG or IPJ blobs anywhere in 7 of
the 22 files: the three 1991 Mount Gordon/Questem files; two Melinda
Downs magnetic files whose AGG siblings from the identical delivery
do have rich REG/IPJ content; and the two derived/inversion-output
databases (East_Isa_VTEM_Inversion.gdb, SAMAGEM_CDI.gdb). Logged
as [LIKELY], not proven: the pattern fits "was this database ever
interactively opened/edited in Oasis montaj" (which is what populates
the registry, per §3.21/§3.22) better than a simple "old files lack
it" theory, since two of the absences are on modern, purely-derived
databases and one is a same-delivery sibling-file split. Recorded as
the most consistent hypothesis given the evidence, not asserted as
settled.
A small, genuinely new, unidentified administrative-blob variant.
DB_AGG_1213.gdb and DB_AGG_1212.gdb each have a handful (6-8) of
administrative blobs tagged with neither REG/IPJ nor the already-
understood all-zero/all-FF empty-placeholder pattern -- a real,
varying 4-byte value instead (L57, abas, etc.). Checked whether
this was a false positive from the line_slot >= 700 administrative-
blob threshold picking up real survey lines instead (it isn't -- this
survey's real lines only go up to line-slot 39, nowhere near 700).
Logged as a genuine, small, unidentified third variant rather than
force-explained.
Write-up. Added NOTES.md §6.9 covering the full pass and both
new findings, updated the "Natural next steps" list (closed the
non-double-array-channel item, added two new items for the REG/IPJ
universality question and the unidentified tag variant), and updated
SPEC.md §5 (VA/array channels) and §9 (REG registry) with the same
findings plus a note in the intro validation-scope paragraph that all
22 files were run together in one combined pass, not just pairwise.
3.24 Reader robustness: fail gracefully instead of hard-crashing¶
A different kind of request from the coordinator -- engineering, not research: the reader should never lose everything it's already decoded to an unhandled exception when it hits a blob, chunk, or record it can't parse. Concretely: a truncated file (cut-off download, or a blob chain that runs past EOF) should make the reader stop cleanly at the point it can no longer make sense of the bytes, hand back everything decoded so far, and warn about the rest -- not throw partway through and discard a chain-walk's earlier valid results.
API shape decision. Went with Python's standard warnings module
(one of the options the coordinator explicitly suggested), via a new
GDBParseWarning class in gdb_reader.py (and a parallel
GRDParseWarning in grd_reader.py), rather than a custom result
object with a partial/errors field. Reasoning: every affected
function already returns a plain list/dict/generator; changing that
return type everywhere would be a bigger, more disruptive API change
than adding a warning alongside an unchanged return value. A caller
that doesn't care can keep using the reader exactly as before; a
caller that does care can wrap calls in warnings.catch_warnings().
lzrw1.py got a narrower, complementary change: introduced a single
LZRW1DecodeError exception to replace what used to be a bare
AssertionError (from assert statements) or uncaught IndexError
(from lzrw1_decompress running off the end of truncated input) --
so gdb_reader.py can catch exactly one well-defined thing instead of
a grab-bag of low-level exception types that happened to leak through.
Went through every function in gdb_reader.py that could crash on
malformed/truncated input, systematically:
- header_fields(): a truncated header (too short for a given int32
field) now sets that field to None and warns, instead of raising
struct.error. This is the root of the dependency graph -- every
other function that consumes header fields (read_channels(),
blob_region_start(), iter_blobs(), find_blob(),
read_blob_values()) needed a None-check added at its own call
site to actually benefit from this rather than just crashing one
level later on None * something.
- read_channels(): bad magic, an unlocatable channel table (caught
find_channel_table()'s existing ValueError, which was already a
controlled exception but previously uncaught by its only real
caller), or a channel table cut off partway through chans_max
records -- all three now warn and return whatever channels were
decoded ([] in the first two cases, a real partial list in the
third) instead of raising or crashing mid-loop.
- iter_blobs(): already naturally "returns partial results" by being
a generator (whatever's been yielded already stays with the caller
when the generator later stops) -- the actual gap was the silence
on why it stopped. Added specific, distinct warnings for: a magic
mismatch after N blobs; a non-positive n_pages; the file ending
mid-header; landing short of true EOF by less than one full header's
worth of leftover bytes (a real, if minor, oddity since every real
file checked in this project ends with an exact match); and a
distinct case found while testing this -- the last blob's own
declared n_pages can imply data extending past true EOF, which
needed its own message rather than reusing the "leftover bytes"
wording (a first draft produced a nonsensical negative byte count
in that message, caught by actually testing against a real truncated
file rather than assuming the logic was right -- see below).
- find_blob(): now degrades the same way if chans_max can't even be
determined (bad magic / truncated header) instead of crashing via a
direct struct.unpack_from call.
- read_blob_values(): a negative row_count (administrative blob --
previously a ValueError, now a warning), an unrecognized channel
type (previously risked a bare KeyError from GS_TYPE_STRUCT[...]
in three separate call sites -- consolidated into one _element_width()
helper that returns None instead), a truncated plain-data read, a
truncated compressed-payload read, an unrecognized chunk subtype, and
a chunk that fails to decompress (zlib.error / the new
lzrw1.LZRW1DecodeError) all now warn and return [] instead of
raising.
- _decode_numeric_or_string(): the one place real partial salvage
was both meaningful and implemented -- if the raw buffer is shorter
than needed for the requested row count (plain/uncompressed data,
file truncated mid-blob), it now decodes as many complete elements
as actually fit and warns about the shortfall, rather than letting
struct.unpack raise and discarding everything. Explicitly did
not attempt the equivalent for a truncated/corrupt compressed
stream (zlib or LZRW1) -- there's no simple way to extract "the first
K decoded values" from a partially-decompressed stream the way there
is for flat data, so those cases warn and return [] entirely; this
asymmetry is called out directly in the docstring rather than left
looking as complete as the plain-data case.
Verification, in two stages. First: re-ran
scripts/full_corpus_sanity_check.py against all 22 real files after
every change and diffed the output against the pre-change run --
zero new warnings fired anywhere, and the decoded output is
identical except for harmless Python set-iteration-order
differences in a couple of projection-name listings. This confirms the
new code paths are inert on real, well-formed data -- they don't
change behavior for the common case, only add a safety net for the
uncommon one. Second: built small, throwaway test files by truncating
real ones at specific points (empty file; header cut to 100 bytes;
truncated mid-symbol-table; truncated mid-blob-header; truncated
mid-blob-chain, past a blob's own declared extent; truncated mid-
compressed-payload) and ran read_channels()/iter_blobs()/
read_blob_values() against each with warnings.catch_warnings(record=
True) to inspect exactly what happened. Every case: no exception
escaped, a real result came back (empty or partial as appropriate),
and the warning text correctly named what happened and where. One real
bug caught and fixed during this testing (not just confirmed clean):
the first version of the "landed short of true EOF" warning could
compute a negative "bytes short of EOF" figure when the last blob's
declared size overshot past the true end of file, producing a
nonsensical message -- fixed by splitting that into its own,
correctly-worded case (off > size) rather than trying to force one
message to cover both directions.
grd_reader.py: a lighter, proportionate version of the same
treatment, since it's a much simpler, single-shot, already-fully-
solved reader with no natural per-blob partial-result boundary the way
.gdb's blob chain has (a .grd file is either a complete grid or
it isn't -- there's no meaningful "first K channels" the way there is
for .gdb's many independent per-line-per-channel blobs). Added
GRDParseWarning; a truncated compressed-block offset/size table, a
truncated or corrupt individual compressed block, and a final decoded-
element-count mismatch against the header's declared shape_e*shape_v
all now warn and return whatever was actually decodable instead of
raising. Verified directly: a real compressed .grd file truncated
partway through its second of five blocks now returns exactly the
16,302 real elements from the one complete block, with a clear
warning, instead of crashing on the truncated second block.
Write-up. Added NOTES.md §6.10 covering the design rationale, a
full list of what changed and why, and the verification methodology;
updated the reader/*.py one-line summaries in §7; updated SPEC.md's
reference-implementation section (§12) with a "Robustness" paragraph.
Also re-ran scripts/lzrw1_full_validation.py (the exhaustive,
non-sampled LZRW1 validator from an earlier round) against a real file
to confirm the LZRW1DecodeError change didn't affect its own,
separate, from-scratch validation logic -- still 1,759/1,759 chunks,
zero failures, unchanged from before.
Session 4 — DB_CHAN_X/DB_CHAN_Y/DB_CHAN_Z channel-role registry keys (2026-09-13)¶
A different kind of question this time: not "what does this format
mean" but a downstream, practical one -- does the format itself help
pick which channels are the X/Y/Z coordinates for a .geoh5-export
feature built on top of this reader, rather than guessing from
channel-naming conventions? That export's x_channel/y_channel/
z_channel parameters currently default to the literal strings
"Easting"/"Northing", which only exactly match a minority of real
files' actual channel names.
4.1 Re-opening NOTES.md §6.8's loose end -- [CONFIRMED] a direct, decodable key -> real-channel-name mapping, on every real file¶
§6.8's "Angle 2" update had already spotted the literal vendor
constant names DB_CHAN_X/DB_CHAN_Y (from the vendor's own
published DB_CHAN_X=0 DB_CHAN_Y=1 DB_CHAN_Z=2 enum, §2) as readable
text in MLMAG.gdb's "REG " registry blobs, "sitting right next to"
coordinate-system metadata -- but stopped there, without pinning down
the exact byte relationship or checking how far it generalizes. This
session did both.
The exact byte structure, confirmed directly. Dumping the raw
bytes around the first DB_CHAN_X occurrence in MLMAG.gdb (real
offset 745564288) shows it sitting inside the same "REG "/"VV "
nested-tag framing already established in §6.8's Angle 1:
... REG \x02\x00\x00\x00\x01\x00\x00\x00\x00\x1a\xcc\xff
VV \x00\x00\x00\x00\x00\xff\xff\xff\x02\x00\x00\x00
DB_CHAN_X\x00x_nad83\x00 ...
i.e. a NUL-terminated key ("DB_CHAN_X") immediately followed by a
second NUL-terminated string -- and that second string, "x_nad83",
is a real, exact channel name in this same file
(GDB(path).channel_names includes it verbatim). Same pattern, same
file, a few hundred bytes later: DB_CHAN_Y\x00y_nad83\x00, and
y_nad83 is likewise real. So this is not merely a constant name
sitting near coordinate metadata, as §6.8 first described it -- it's
a genuine, directly-decodable key -> channel-name mapping, naming
exactly which channel plays the X (or Y, or Z) role, with no
naming-convention guessing needed at all.
Corpus-wide check, all 22 real files, all 3 agencies: universal for
X/Y, real but less common for Z. Scanned every real .gdb file in
samples/ for DB_CHAN_X\0/DB_CHAN_Y\0/DB_CHAN_Z\0 and decoded
the NUL-terminated value immediately following each occurrence:
| Key | Present | Value verified against the file's real channel table |
|---|---|---|
DB_CHAN_X |
22/22 (100%) | every file |
DB_CHAN_Y |
22/22 (100%) | every file (see staleness caveat below) |
DB_CHAN_Z |
5/22 (23%) | every file |
[CONFIRMED] on real production data across USGS, GSQ, and Ontario deliveries alike -- not an artifact of the one Ontario file that first turned it up.
A real, meaningful "no channel assigned" value, distinct from
absence. Three files (AG106386_Northern Georgetown_Conductivity.gdb,
DB_EM_293.gdb, East_Isa_VTEM_Inversion.gdb) have a DB_CHAN_Z key
whose value is a single literal space character (" ") rather than a
channel name -- confirmed by direct byte inspection, not a parsing
artifact or a truncated read. Read as Oasis montaj's own explicit
"this role has no channel" placeholder, distinct from the key not
existing at all.
A real complication: stale, superseded copies of the same key can coexist in one file, and neither "first" nor "last" wins reliably. Two files have multiple, differing occurrences of the same key:
SAMAGEM_CDI.gdb(this project's largest real file, 1.93GB,DB_COMP_NONE): fourDB_CHAN_Yoccurrences -- three read"Yg"(not a real channel in this file's current channel table) at offsets 205184/207232/1855099264, and one reads"y_NAD83"(a real channel) at offset 1855108480, the very last one in the file, right next to aDB_CHAN_X\0x_NAD83\0at offset 1855108224. Here the last occurrence is the correct one.DB_Mag_833.gdb: twoDB_CHAN_Yoccurrences --"Northing_AGD66"(a real channel) at offset 980352, and"Y"(not a real channel) at offset 4048256. Here the first occurrence is the correct one -- the exact opposite of the previous case.
Both are consistent with this format's general append-only,
never-in-place-edited blob storage model (already established
elsewhere, e.g. §6.6b's blob-chain findings): re-registering a file's
X/Y channels in Oasis montaj evidently appends a fresh registry entry
rather than overwriting the old one in place, and there's no reliable
positional rule (favoring first or last) for picking the live one
after the fact. The rule that resolves both real cases correctly:
validate each candidate value against the file's own real channel
table (read_channels()), and keep whichever occurrence(s) actually
name a real channel -- a direct cross-check against data this reader
already has to parse anyway, not a guess. No file in this corpus had
two different candidate values that both matched real channels, so
genuine remaining ambiguity (a role reassigned to a different,
still-live channel) hasn't been observed -- only reasoned about as a
theoretical edge case this rule alone wouldn't resolve.
Why this matters, concretely. The .geoh5-export feature's own
coordinate-channel defaults currently hardcode the literal names
"Easting"/"Northing", which (checked directly against this same
22-file corpus) exactly match only 3 of 22 real files -- every other
file uses a different real convention (EASTING/NORTHING,
MGA_East/MGA_North, x_nad83/y_nad83, UTMX/UTMY, plain
x/y, ...). This registry mechanism, once decoded, resolves the
correct channel on every single file in the corpus (100% for X/Y) --
a dramatically better default than any hardcoded name convention could
be, with no guessing involved at all.
What's still open: the exact binary field boundaries around the
key/value pair (same honest gap as the rest of §6.7/§6.8's "REG "
framing -- found by searching for readable text within an
already-tag-framed region, not by parsing a byte-exact record layout);
whether a file can have genuine, unresolvable ambiguity (two differing
values that both match real, currently-live channels) -- not observed
in this 22-file corpus, but not proven impossible either.
Write-up. Added NOTES.md §6.8b with the same findings, framed as
a direct continuation of §6.8 rather than by session. Not yet wired
into a reader function -- this session was scoped to confirming the
finding and its reliability, not implementing a decoder or changing
to_geoh5's defaults.
Session 5 -- a user-reported silent truncation: DB_COMP_SPEED blobs are chains of chunks (2026-09-23)¶
A user reported a problem with a real file and supplied it for
debugging, on the condition that nothing identifying it be recorded. It
is therefore described here only as "the supplied file": several hundred
MB, DB_COMP_SPEED, a handful of lines, a few dozen channels.
What was asked. Work out what goes wrong reading it.
What was tried, in order.
- Opened it: no exception, no warning. Read every populated channel on every line: still no exception -- but every numeric channel had exactly 2046 rows and the 255-wide string channels exactly 64, from a file whose real-line blobs held hundreds of MB. That mismatch (a few MB decoded out of several hundred) was the first sign the reader was quietly under-reading rather than failing.
- Ruled out the container:
iter_blobswalked the chain to exactly the file size, so no blobs were being missed. The lines and populated channels were real; the contents of each blob were short. - Noticed 2046 x 8 = 16368 bytes, the ~16KB "largest real chunk" that
rust/src/lib.rs's module docs already mentioned -- one chunk's worth. Readread_blob_values: its DB_COMP_SPEED branch parsed exactly one chunk (parse_chunk_header(raw_span, 0)) and stopped. - Confirmed on raw bytes: a blob's span held one magic occurrence, but
the first chunk ended after ~2KB of a ~350-page blob. At
16 + chunk_lengthsat a bare 12-byte sub-header, not a magic -- the samef0 3f 00 00(16368) + plausible length +0xF4E5D6C7marker -- and the int32 16368 recurred at exactly the offsets a chain of such chunks predicts. (NOTES.md section 6.6e.) - Read the 56-byte blob header for a way to know where the chain ends:
+24matched the sum of every chunk'sdecompressed_lengthon all of the file's real-line blobs;+28matched16 + sum(chunk_length);+48matched the row count. Also found the bytes after the last chunk are non-zero padding, so "stop at zeros" is not an option. - Checked whether this was specific to the supplied file: it is not.
Ran the same header-vs-chain comparison over this project's own
corpus: 1,656 of 7,015 Speed blobs are multi-chunk (e.g. one in
DB_EM_293.gdbwith+24= 16480, first chunk 16368), all of which the reader had been silently truncating.DB_COMP_SIZEwas checked too and is unaffected (a single zlib stream,+24-exact on 3,414 of 3,414 blobs). -
Decoded every chunk in the supplied file after the fix: numeric channels on a line now agree on row count (hundreds of thousands, up from 2046), string channels match them, values are continuous across the 2046-row seams. One channel on one line is genuinely shorter -- its own header
+48row count says so -- which is what the file declares, not a decode error. -
Checked the fix against independent ground truth. The user also supplied companion spreadsheet exports of the same survey (one workbook per lines' worth of data). Aligning spreadsheet rows to decoded rows by the ID column: every spreadsheet row was found in a decoded line (100% for two workbooks, all but 20 rows for the third), and for the full-length lines 12-14 of the 15 numeric columns match a decoded channel exactly on every row -- including all the rows past the first 2046, at the chunk seams and at the tail, which the old reader never returned. That is the first check in this project of the chunk-chain decode against data that did not come from the file.
- The columns that did not match turned out to be a second, separate
problem, not a decode error. Three spreadsheet columns per workbook
(one in the third) had no exact counterpart -- but their values were
exactly the same multiset as a decoded channel's, in a different row
order (smooth in the spreadsheet, scrambled against the rest of the
line in the decoded data). Decoding the first blob for that
(line, channel) directly reproduced the spreadsheet row for row; the
reader (
GDB._ensure_blob_index, "last blob wins") had returned a later blob for the same (line, channel). The file holds 30 of 110 real (line, channel) pairs twice -- two blobs, sameblob_index, same values in two different row orders. Comparing both copies with the spreadsheets showed neither "first wins" nor "last wins" is right: the last copy is current for two pairs on one line, the first for six pairs on the other lines. Nothing in the two copies' blob headers distinguishes them (identical timestamp, totals and reserved fields; only the compressed size differs), and the metadata region before the first blob holds no directory of blob page numbers or offsets (searched for every blob's page number and byte offset: only chance hits). NOTES.md section 6.6f. Not fixed here -- how the format marks the current copy is still unknown. (Superseded: Session 6 found it -- a 6-byte-stride directory at offset 280 that this search's encodings missed; NOTES.md section 6.1c.) -
Not specific to the supplied file: 2 of the 22 corpus files also have duplicated (line, channel) blobs (345 of 116,683 pairs), so the reader's "last wins" has been an untested assumption there too.
-
Tried to find how the format marks the current copy -- nothing found. Diffed every header byte and the chunk sub-header of 13 duplicated pairs' copies (only sizes differ; every timestamp is
INT_MIN); looked for a directory of blob locations in the metadata region, in all 66 administrative blobs, in the line/channel records and (a whole-file search) for 26 blobs' (start, size) pairs in 9 encodings -- none; scored every simple size/position rule against 13 labelled pairs -- none fits (best 10 of 13). What did come out: the allocation slack pattern (whole-extent re-use of freed space, so a newer blob can precede the older one), and that the corpus's own duplicates look like the plain "append the rewrite, abandon the old blob" case (206 pairs of one revised channel in one file, headers byte-identical, values revised not reordered; 139 pairs with identical values in another) -- no independent ground truth for either, so which copy is current is still an inference there. NOTES.md section 6.6f. Still [UNKNOWN] how the current copy is marked; possibly it is not persisted at all. Also checked all 22 corpus files: no pointer table in any; blob timestamps set in only 2 files, neither with a duplicated data blob;n_pages!=n_pages_dupin 6 files (extent length vs pages in use, [LIKELY]); header word 116 non-zero in two files, meaning unknown. 12. Looked for processing steps that might explain the reordered copies, and whether the format records "channel math" at all. It does -- spec section 9's registry is per-channel, in the supplied file as in the corpus (FORMULA= the expression,MAKER= tool + saved parameters, including tool records for channel slots with no channel-table entry, and large registry-format objects whose contents are undecoded). Tools recorded: channel math, 1-D FFT filter, non-linear filter, polygon mask, grid sampler. Recovering the row mapping between each duplicated channel's two copies showed it is one operation shared by every duplicated channel on a line, and that the stale order is exactly sorted by the X coordinate (registry-confirmed slot 2) on two lines, with the current order sorted by ID/date/time -- i.e. temporary spatially-sorted copies of a subset of channels. Which tool left them is a [GUESS]; no dates exist in the administrative blobs to order the steps. NOTES.md section 6.6f.
What it changed. NOTES.md section 6.6e (new) and docs/spec.md
sections 7.3-7.5: the chain-of-chunks framing, and +24/+28/+48
promoted from [LIKELY] "first-chunk preview" to [CONFIRMED] with their
real meaning. Reader: decode_speed_blob in pygdb/lzrw1.py and (kept
in step) rust/src/lib.rs, used by read_blob_values.
Why it slipped through. Every earlier ground-truth check compared
leading values (which the first chunk gets right) and small blobs;
nothing compared a decoded length with an independent row count. +48
is exactly such a count, and is not yet used as a cross-check.
Session 6 -- reading the pre-blob region against the vendor's constants (2026-09-26)¶
Prompted by: "there's still a lot left to determine about the pre-blob region -- did you look through there, and are there any constants in the open-source repo that might help identify anything?"
Honest starting point. Session 5's search of the pre-blob region only looked
for a directory of blob locations (none) and noted the empty and symbol-table
pages; nearly every header word was still [UNKNOWN] in the spec. The
vendor-constants block already recorded in NOTES.md section 2 had never been used as a
key for the header.
- Read the header words against the vendor's
DB_INFO_*order andDB_SYMB_*kinds. Words 84/88/92/96 turned out to be the four symbol-table capacities inDB_SYMB_*order (92 and 96 equal the knownchans_max/users_max); 72/76/80/64 are their running totals; 48/52/56/60/44 partition theblob_indexspace; 104 is the end of the symbol-table region and 108/112 follow from it. Every relation holds on 23 of 23 files (NOTES.md section 6.1b). - Laid out the region with that arithmetic: header, front block, blob-symbol
table, line table at
chan_table - 24 - lines_max x 128(exact; identical lines in identical order in 23 of 23 files, four with the documented index shift of -1), channel table, user table, 8 bytes, padding. The heuristic line- table search is no longer needed to find the table. - Followed the odd header word (32: 100 in most files = the documented
GXDBdefaultcache=100, much larger in others). Between the two files with identical capacities the region after the header differs by exactly 6 bytes per unit. First reading (wrong stride, corrected in step 8): repeated 12-byte empty entries about half that many, and 12 used entries in the supplied file; the two files with no entries at all were the only two with header word 116 non-zero. - Decoded the non-empty entries in the supplied file: a blob's relative start page
and
n_pageson every one -- including one that is the older copy of a duplicated pair. First conclusion (wrong, corrected in step 8): "a blob cache table, not a directory of all blobs". - Dead ends recorded: no page handle or blob position in any of the 350 blob-symbol records; a naive "line table is sparse symbol records" check for the blob-symbol table classified only 12 of 23 files cleanly (the populated ones use a layout not decoded), so that table's position is stated as [LIKELY].
- Asked, for the next request ("account for every non-zero byte before the first block, label it by what it is and its range, then list what is left"), for a byte-accurate accounting tool: it labels every non-zero byte of the pre-blob region as decoded / likely / structure-only / unlabelled and reports the rest. Its first version could not place 6-byte-stride slots under the 12-byte reading, which is what exposed the stride error.
- Re-read the front block at a 6-byte stride starting at offset 280, indexed by
blob_index: an entry(uint32 word, uint16 n_pages)with flag0x8and a relative start page landed exactly on a blob header whoseblob_indexwas the slot, for all 104 listed (line, channel) entries of the supplied file; in the corpus 116,287 of 116,287 across 23 files, zero failures (NOTES.md section 6.1c). - Corrected step ¾. The "cache" is the tail of the same array (
word 60 .. word 44), and the array is the directory of live blobs, not a cache alone. For issue #2: in the corpus all 345 duplicated pairs are listed at the last copy; in the supplied file 19 of 27 listed duplicated pairs are listed at the first copy, and all 13 independently labelled pairs agree with the directory. Word 116 is non-zero in the two files whose cache slots are full (99/100 and 100/100), not "empty" as first written. - A zero directory slot marks a blob the file does not list as live. In the corpus
that is 582 blobs in
SAMAGEM_CDI, all of them two whole channels of 291 lines each, no duplicates, both still in the channel table (CVG_GSCLevel,mag_gsclevel); the supplied file has six such blobs in two channels. Why is unknown -- an earlier remark that they are "probably deleted channels" is not established and is not written into the spec. - Census of what is still unlabelled after all of the above (NOTES.md section
6.1c): about 496 KB of non-zero bytes across the 22 corpus files, in fixed
record fields of the blob-symbol, line and channel tables plus a 24-byte gap,
all
[UNKNOWN]. Many are float dummy values (float32/float64-1e32). The directory's last five slots overlap the nominal start of the blob-symbol table by 32 bytes in every file, so that table's start is[LIKELY]only. - Wired into the reader (behaviour and validation in section 6.1c and spec 2.2):
read_blob_directory, strict per-entry validation, directory-selected live copy, skip-unlisted with an opt-in, exact line-table position with the heuristic as fallback; theduplicate_blobspolicies were removed. Checked against the old reader on the corpus: identical line names/indices in 22 of 22 and identical blob selection in 21 of 22 (the 22nd drops exactly the 582 unlisted).
What it changes. docs/spec.md section 2 (header table), new sections 2.1
(layout) and 2.2 (the blob directory), 6.2 (there is an index of live blobs), and
the reader. For issue #2 it settles which copy is current wherever a file has a
directory (all 23 here); the earlier row_order heuristic is no longer needed and
was removed. Open: word 116, the 0x4 flag bit, why blobs are unlisted, and every
field in the section 6.1c census.
Session 7 -- decoding the "REG "/"VV" administrative-blob framing (2026-09-28)¶
Prompted by: "are there any parts of our spec that have not landed into the python package?" (identified section 9's registry content as description-only, no decoder), then "are the registry contents you've decoded structured in any way?", then "can you decode it?"
- Dumped a real
REG-tagged administrative blob byte-for-byte (Magnetic_Data.gdb, blob_index 50123, offset 115064832) and read it word by word. The long-standing[UNKNOWN]constant4670802(spec section 6.4, "a specific non-GS_*constant... instead of a valid type code") is exactly the ASCII bytes52 45 47 00("REG\0") read as a little-endian int32. It was never an opaque sentinel -- the blob-header field a real data blob uses for itsGS_*type code (relative +44) is, on an administrative blob, the first 4 bytes of the object's own 3-letter name. The same holds forIPJblobs (49 50 4a 00). - Confirmed a fixed 128-byte preamble on every
REGblob: the ordinary 48-byte blob header (name in place of the type code, as above; timestamp always the0x80000000unset sentinel; a "kind" field at +20 always100, distinct from the200/202seen on real data blobs), then 48 more bytes ending in a repeating constantff 00 e1 1e, then a 16-byte block (separator00 1a cc ff+ FourCC tag"REG "+ two int32 counts2,1), then another 16-byte block for a nested"VV "tag (Geosoft's own vector-value object, matching the framing sketch already in section 6.8). Exact on 1,033 of 1,033 real REG blobs checked across 5 files, all 3 agencies (Magnetic_Data.gdb466/466,Radiometric_Data.gdb356/356,MLMAG.gdb65/65,MLGRAV.gdb63/63,AG106386...Conductivity.gdb83/83). - Decoded the VV object's content past that 128-byte mark, and found it is not one thing:
- A flat cached numeric array. One VV (blob 50123 above) is a
contiguous run of 112 real float64 values, smoothly varying (279.3 down
to 178.0), filling the rest of the page. Real, decodable data -- but
what it is remains open: it isn't a copy of the real per-line data for
the same channel_slot (checked directly, channel 23 is
diurnaly_cor_mag, whose real per-line values are ~47,651, a different range entirely). - A short flat
KEY\0value\0pair per ~256-byte slot. Another VV (blob at 229258240) holdsFORMULA\0time(hh,mm,ss)\0, thenLABEL\0(empty value), thenUNITS\0(empty value), each starting at(offset - 128) mod 256 == 0relative to the blob. This period holds for the majority (70-90%) of everyFORMULA/UNITS/LABEL/CLASSoccurrence checked across 4 files -- but not all of them (see next). - A second level of the same recursive framing. A third VV (blob at
229257216, one page before #2 above) recurses: at +384 a miniature copy
of the same shape (
1, the constant0x0ff000ff, a length,0) then1,"MAKER\0", theff 00 e1 1econstant, a length,0,1, then00 1a cc ff+"MAKE"+ count1+ length80, then 80 bytes of the literal tool identifier string ("geogxnet.dll(Geosoft.GX.MathExpressionBuilder...)") -- the same name-header -> marker -> separator+tag shape as the outerREG/VVpair, one level deeper. This is genuinely[UNKNOWN]which of the two content kinds (flat slots vs. recursive object) a given VV holds, or what decides it. - The GX-tool parameter block itself is plain text, not further
TLV-tagged. Right after the
"Channel Math Expression Builder"display name in the blob above: a UTF-8 BOM (ef bb bf), then CRLF-separatedKEY.SUBKEY="value"lines --MATHEXPRESSIONBUILDER.CHANNELINPUTBOX="ch_9=comp_mag - ch_8;ch_9=ch_9 + 48066.0;"among them, the same real formula already known from section 6.8 -- ending with a0x1A(DOS/ASCII SUB, a classic text-file EOF marker) byte. [CONFIRMED] on this one instance; not yet checked elsewhere. - Not chased further this round (recorded honestly, not guessed at): what
selects flat-slot vs. recursive-object content for a VV; the meaning of
the
0x0ff000ff/ff 00 e1 1e/00 1a cc ffconstants; why some 256-byte slots hold a real key and others are empty placeholders (open since section 6.8); the complete list of top-level tags beyondREG/IPJ/VV/MAKER/MAKE/CLASS. - Asked whether the "kind" field (+20, always
100in every example dumped so far) correlates with the three content forms. It does not -- it is100on 1,942 of 1,942 administrative blobs across the whole real corpus, every real tag seen (REG\01,758,IPJ\063,EXT\036,META\04). It marks "administrative blob" as a class, not which registry object a blob holds. - The length-like fields at +24/+32/+64 (step 2) do correlate with
content kind, cross-tabulated against a content classifier (empty /
numeric array / flat key-value slots / nested sub-object) on 4 files:
empty, numeric-array and other/short all sit at a fixed
104baseline; nested sits in a tight per-file cluster above it (283-285USGS,391-395GSQ); flat key/value is wide and tracks actual string length.+24is therefore a real length field -- "declared payload before any undeclared tail" -- not a discriminator on its own, and a numeric array's real length isn't in it at all. - Chasing step 6's tag census turned up two more real variants, one a
false positive of the census method itself.
"LINE"is not a real object tag -- all 4 hits (1991 Melinda Downs GSQ files) are a blob with noREG/VVwrapper at all: after the ordinary 24-byte prefix and a length field at +24 that exactly matches the text length in every instance (916/691/1317/461 bytes), raw CRLF text begins immediately at +32 -- a literal Oasis montaj "OASIS VIEW" saved session/settings block ([OASIS VIEW]\r\n\r\nLINE L570300\r\nCHANNEL EASTING\r\n..., a per-channel display-profile list). The word "LINE" in that text simply landed on the same +44 offset a real object's name occupies, in these 4 instances, which is how the tag census in step 6 miscounted it."META \0", unlike "LINE", is real -- 4 of 4 instances the same shape, 3 files, all 1991 GSQ TEM/EM surveys. Same 128-byte preamble asREG, but the first nested tag at +96 is"ATEM"(plausibly Airborne TEM), and a short distance further in, the 16-byte page-primitive magic (section 6.5/7.1) withsubtype=2-- a genuine embedded zlib stream, confirmed on 4/4. Not decompressed or examined further. - Tried the same cross-tabulation against the two VV fields not yet
checked, +116/+120 (never vary, no signal) and +124. +124 is the count
of distinct keys in a flat key/value VV -- 243 of 245 clean instances
(first 256-byte slot itself empty or a real key) match exactly across 5
files; the 2 exceptions (
MLGRAV.gdb) are the same real object withUNITS\0duplicated across two slots and+124=1, i.e. it counts distinct names, and a duplicate collapses to one -- the same append-only stale-copy phenomenon already known for data blobs (section 6.6f). A real subset (57/102 inMagnetic_Data.gdb) doesn't fit this model at all: a non-key value (a bare int32,83seen) sits in the first slot, and +124 is 0 regardless of how many real keyed slots follow -- left open. - Looked closer at that leading value: it's a single byte at blob-relative
+132 (the preceding 4 bytes are always 0), not an int32. Across all 80
such blobs in
Magnetic_Data.gdb:83(ASCII'S') on 50, but 21 further distinct values seen otherwise, each mostly once -- a dominant common value with real per-instance variation, so genuine per-blob data, not a fixed marker, but nothing decoded so far says what it holds. - Decompressed the
"META"/"ATEM"zlib stream found in step 8 (had a trivial bug first -- compared a 4-byte tag slice to a 5-byte literal, silently matching nothing; fixed and reran). All 4 instances decompress cleanly with plainzlib(11,094-12,211 bytes), each distinct (4 different SHA-1 hashes, including the two within one file -- presumably a stale/current pair). It is not survey data. Its readable strings are Geosoft's own internal class/type vocabulary --"IPJ Class","ITR Class","DOCU Class","META Class"(explaining the outer tag),"PLY Class","PIC Class","TPAT Class", primitive types (Bytes,String,Object,Enum,Bool,Angle,Time,Date,Data,Picture), attribute names (FixedSize,MaxSize,ByteOrder,MinValue,MaxValue,EnumValue,Visible,Editable,FlatName), headed by"Geosoft","Core","Types","Objects"-- the same category of thing as the already-documented bundled projection dictionary (section 6.7), a generic reference catalog the software embeds, not something specific to this survey. The record framing inside it (repeating0x02+ a one-letter kind code + several int32 fields + an optional name) looks real and regular but wasn't decoded further -- low priority, since it describes the software, not the data. - Asked directly what selects flat-slot vs. nested
VVcontent. Cross-tabulated every fixed field+0..+124against content kind on 3 files: nothing but+24/+32/+64(already known to correlate) differs at all between kinds -- there is no separate flag. But+24alone works as a practical discriminator: on every file withnestedinstances, its value sits in a narrow band (one value or a tight cluster per file --283,285,391-395) that zeroflat_kvinstances ever land inside, across all 5 files checked --flat_kv's own values are either the104baseline or jump straight past the nested band. Recorded as an empirical gap that holds everywhere checked, not a proven encoding rule. - Pushed on the bare-placeholder question (open since section 6.8), broken
down by key name instead of treated as one phenomenon:
CLASSis always empty (0 of 216 real instances, all 5 files) andFORMULAis always populated (0 of 69 empty, all 5 files) -- both fully deterministic, not "sometimes". OnlyLABEL/UNITSgenuinely vary. Tested four hypotheses for what decides it on all 5 files: whether the channel has real per-line data anywhere, its dtype (string/numeric), whether it's an array channel, its display format code -- every combination has real counts in both directions on every file, none ruled in. Still open, four real candidates ruled out. - Tried the remaining constants. Census over the whole corpus (not just
the handful of files checked before) confirmed
+28/+60are fixed on 1,861 of 1,861 real administrative blobs -- never a third value. Both turned out to be two complementary byte pairs:(0xFF, 0x00)then(X, ~X),X = 0xF0at+28and0xE1at+60-- verified by XOR (byte0^byte1 = byte2^byte3 = 0xFFon both). Checked the separator (00 1a cc ff) against the same shape: it doesn't fit (XORs0x1a/0x33, not0xFF) -- a distinct, unrelated constant. The bit structure is real and complete; what0xF0/0xE1themselves mean is not.
What it changes. docs/spec.md sections 6.4, 9 and 11 (the 4670802
constant is no longer [UNKNOWN]; the framing has a real, if partial,
decoded structure now, superseding "no decoded byte-exact structure, just
plain string search"). No code changed -- this is investigation only, kept
out of pygdb/registry.py pending a decision on whether a real decoder is
worth building on top of a structure this partially understood.
Session 8 -- the IPJ coordinate-system record's fixed byte offsets (2026-09-28)¶
Prompted by: "Can we push on the Coordinate system meta data?" -- a direct
follow-up once Session 7's "REG " framing made it worth checking whether
IPJ blobs share the same preamble.
- Dumped a real
IPJ-tagged blob (AG106386_...Conductivity.gdb) and confirmed it shares Session 7's exact 128-byte preamble (the0xff 0x00 0xe1 0x1econstant, the0x00 0x1a 0xcc 0xffseparator, at the identical offsets). A second nested tag was found at +112, after the already-known" JPI"name marker at +96 -- a 4-byte FourCC abbreviation of the grid system (" UTM","MGA "), not decoded further. - Field-by-field dumped the content past +128 and searched every real IPJ blob (>=640 bytes) across 5 files/3 agencies for each file's own independently-known real geodetic constants' exact float64 bit patterns. All landed at the same absolute byte offset in every instance: datum name +180, ellipsoid name +244, semi-major axis +308, eccentricity +316, central meridian +596, scale +620, false easting +628, false northing +636 -- 30 of 30 real instances, GSQ/Ontario/USGS all agreeing, replacing the previous "consecutive small byte deltas" description with exact offsets (NOTES.md section 6.7b).
- Two more float64 slots (+604, +612) sit between central meridian and
scale; every one of the 30 instances reads exactly the vendor's
rDUMMYsentinel (-1e32) there -- real fields, never seen populated. - What first looked like the offset hypothesis failing on
MLMAG.gdb's andMagnetic_Data.gdb's minor IPJ objects (ellipsoid-only, no projection) turned out to be the same dummy-value convention correctly marking "no projection defined" at +596 onward -- a confirmation, not a gap. - The datum-transformation name (
"GDA94 to WGS 84 (1)"etc.) at first appeared to sit at +332 on 11 of 11 GSQ instances but not on the Ontario/USGS files checked -- looked like a real agency difference and was briefly written up as one. - Reinforced (not newly found) the existing note on serialized in-memory pointers: +136..+176 reads as classic Windows x64 pointer shapes (0x00007ffd...) on more than one real instance dumped this round.
- Kept pushing (user: "sure keep pushing"). Checked +604/+612 across the whole 22-file corpus, not just 5 files: 63 of 63 real instances, still always the dummy -- a stronger negative result, not a new one.
- Caught and fixed a real mistake from step 5. Dumped the raw bytes at
+332 directly on the exact Ontario/USGS blobs step 5 said didn't have it
there -- they did:
"NAD83 to WGS 84 (1)"right at +332, byte for byte. Step 5's test had aggregated hits from 3 files into one counter and silently lost the Ontario/USGS ones, a script bug, not a format difference. Re-verified properly (explicit per-blob check, whole corpus): 36 of 36 real instances that define a transform have it at +332, all 3 agencies, zero exceptions. Corrected NOTES.md section 6.7b and docs/spec.md section 8 rather than leaving the wrong claim standing.
What it changes. docs/spec.md section 8: the vague "consecutive small
byte deltas" description is replaced with an exact offset table, all
[CONFIRMED] (including +332, corpus-wide, after the step-8 correction)
except +604/+612 (structure confirmed on the whole corpus, meaning still
open). No code changed -- investigation only, same
disposition as section 6.8c: not wired into pygdb/registry.py pending a
decision on whether a decoder is worth building on it.
Session 9 -- decoding the line and user records overnight (2026-09-28/29)¶
Prompted by: "Lets try to decode the line and user records, go ahead and work on both of these for a while overnight." Unattended follow-up work; no code changed, investigation and documentation only, same discipline as Sessions 7-8 before anything was wired into the reader.
- Byte-censused every int32 field of every real line record (
db.lines, the exact-offset table from the header-geometry work) across all 22 real files, 5,003 real lines total. Six fields turned out to hold real, confirmable content; two more are confirmed-reserved zero; two remain genuinely unresolved. +124is the line's own integer line number. Cross-checked against every line name ending in a number: 4,994 of 4,994 exact matches, zero exceptions. For 9 real lines with a.1/.2repeat suffix,+124holds the integer part only. Compared one such line byte-for-byte against its immediate non-repeat neighbours ("L1004.1"vs."L1003"/"L1005","L1263.1"vs."L1262"/"L1264",MLGRAV.gdb): the only differences anywhere in the 128 bytes are the name text and+124itself -- the.1suffix has no separate storage at all, confirmed negative, not just unfound.+0correlates with the line-name prefix letter:2on 225 of 234 real"T"(tie-line) names,0on all 4,769"L"names. The 9Texceptions are, in every one of the 9 files with anyTlines, exactly that file's own single lowest-numberedTline (T101in bothMLMAG.gdb/MLGRAV.gdb,T7000in both USGS files) -- checked directly, 9 of 9 -- plausibly a reference/base tie line.+8..+27(20 bytes) is always the identical byte pattern when populated (4,985 of 5,003 real lines, every file) -- decodes as the vendor's-1.0e32/+1.0e32dummy sentinels. Real content for this field has never been seen; same "confirmed structure, dummy-only" shape as the IPJ record's+604/+612(section 6.7b).+116..+123decodes as a real float64 -- a decimal year. Values cluster in 2003-2025, and within one file cluster tightly around one date rather than spanning a real survey's weeks-long acquisition window (Magnetic_Data.gdb: 631 of 631 lines read the identical2020.2295= 2020-03-25) -- read as a per-file creation/save timestamp, not a per-line flight date. Present on 18 of 22 files; absent on the 4 oldest (1991 GSQ) files in the corpus.+96..+107and+112are always exactly zero -- 5,003 of 5,003, confirmed reserved/unused.- Moved to the user table. A first, loose census (any NUL-terminated-looking
name) found several "extra" users in 3 files -- all false positives,
leftover embedded-blob bytes landing in unused table capacity by chance
(the same phenomenon already documented for the channel table's own
false positives). Re-ran with the project's existing
_read_nameclean check: every one of the 22 real files has exactly one real user, slot 0, the superuser -- no real second user ever found in this corpus. - The name field is only 32 bytes wide for a user record (not 64 like channel/line records) -- inferred from real content resuming cleanly at +40, not directly proven (no real user name here is long enough to test the boundary).
+40..+71(32 bytes) is a UTF-16LE fragment of the file's own creation/ save path. Decoded text matches the real file's own name on 13 of 22 files ("AGG_1212.gdb","Gordon_1003.gdb","\Device\..."on the two USGS files, etc.) -- clean and NUL-terminated when the real path is short enough, silently truncated (losing the terminator, sometimes the very last character) when it isn't, with leftover non-zeroed bytes after a real terminator -- the same "unused capacity isn't reliably zeroed" behavior already seen for the channel table and the REG registry's dirty-slot-0 case, now confirmed a third time. This generalizes docs/spec.md section 3.3's single earlier sighting ("observed once") to 13 of 22 real files with a precise, confirmed offset.+72..+79is byte-for-byte the same per-file timestamp as line+116-- checked directly onMagnetic_Data.gdbandRadiometric_Data.gdb, exact match both times. Independent cross-reference between two different table types, not just a similar magnitude.+84is always exactly131072and+124is always exactly-1on every one of the 22 real superuser records -- confirmed values, meanings not established (no second, non-super user exists in this corpus to compare against).- Went back to the two line-table loose ends. The
.1/.2repeat suffix has no separate storage anywhere: comparedMLGRAV.gdb's"L1004.1"byte-for-byte against its immediate non-repeat neighbours"L1003"/"L1005"(and"L1263.1"against"L1262"/"L1264") -- the only differences anywhere in the 128 bytes are the name text and+124itself. A confirmed negative, not just "not found yet". - The 9
+0exceptions are each file's own single lowest-numberedTline -- checked directly on all 9 files with anyTlines, 9 of 9 (T101in both Ontario files,T7000in both USGS files). Plausibly a reference/base tie line. +4is plausibly a flight number. Mostly zero; on the one file with enough real variation to test (AG106386_...Conductivity.gdb, 112 real lines, 58 distinct values 0-69), 48 of 58 values are shared by exactly two lines, and every such pair is two directly-adjacent line numbers (ten apart, this file's own real numbering increment) -- the shape a flight covering an out-and-back line pair would have. No independent flight-count record was available to confirm it against.
What it changes. docs/spec.md sections 3.2 and 3.3: both gain a real
offset table in place of "everything else: not decoded". provenance/
notes.md gains section 6.2c (user table) and 6.3b (line table). No code
changed -- same as Sessions 7-8, this is investigation and documentation
only, left for a follow-up decision on whether to build real
LineRecord/UserRecord field readers on top of it.
Session 10 -- three follow-up leads (2026-09-29)¶
Prompted by: "do those two you suggested. then see if you can do the flat vs nested selector" -- the three items flagged as most promising when asked "any more leads left to follow up on?".
- Line record
+12..+19(the middle 8 bytes of the+8..+27dummy block), decoded precisely. Reads as the closest float64 to a clean, round-9x10^31-- not-1e32itself, and not previously catalogued as a vendor dummy value anywhere in this project's constant research. Suspiciously round to be coincidental, but not established as a real third sentinel either; recorded honestly as unexplained-but-noteworthy rather than claimed. - The REG dirty-slot0 byte correlates with the specific set of keys
later in the object, corpus-wide -- real for two values, noise for
the rest. Grouped every dirty-slot0 object's byte against its exact
key set:
83/'S'maps to(LABEL, UNITS)or(FORMULA, LABEL, UNITS)on 75 of 80 instances;68/'D'maps, 10 of 10, to objects containing at least_PJ_ELLIPSOID/_PJ_IPJ/_PJ_NAME(the projection-serialization keys, section 6.7's textualIPJcopy). Both read as real, plausible mnemonics (Settings / Datum).72/'H'also repeats cleanly (2 of 2,DB_CHAN_Y/DB_CHAN_Z) with no obvious mnemonic. The remaining ~90 distinct byte values are each seen on exactly one object in the whole 22-file corpus, spanning the full 0-255 range near-uniformly -- the signature of uninitialized memory, not a per-tool marker, matching "dirty slot 0" already meaning something here didn't get written normally. - Two more hypotheses for the flat-vs-nested
VVselector, both refuted. Neither the administrativeline_slotnamespace value nor which real channel the entry is about determines content kind: every adminline_slotchecked holds a mix of kinds, and 15 of 50 real channels onMagnetic_Data.gdbhave REG entries of more than one non-empty kind across their different administrative blobs. With the original fixed-field census, the+24-magnitude proxy (section 6.8c) remains the only signal found; content-probing is genuinely necessary.
What it changes. docs/provenance/notes.md section 6.8c gains the
dirty-slot0 key-set table and the two refuted selector hypotheses; the
line-record +8..+27 note (section 6.3b) is not changed in status (still
"real content never observed"), just given a more precise decode of its
middle bytes. No code changed -- investigation only.
Session 11 -- symbol records begin at their name; channel display width/decimals; REG objects are named by symbol handle (2026-09-29)¶
Prompted by: "I was working on this with sonnet, see if you can identify any
more angles to try to solve any undecoded items." Scripts in the session
scratchpad (angles1.py..angles11.py); every number below is corpus-wide
over the 22 real files unless a file is named.
- New oracle: the ASEG-GDF2
.dfnsidecar forAG106386. It had only been used for the VA/array finding (section 6.2b), never for display settings. Each field carries a Fortran-styleFw.d/Iwformat. Against the channel record:+96equals the.dfndecimal count on 37 of 37 channels the.dfndescribes (including both 30-element array channels);+94equals the field width on 34 of 35 scalar channels (GA_project_number:+94=6,.dfnI4), and differs on both array channels (LEI_Conductivity8 vs 10,LEI_Depth10 vs 12). The USGS CSVs were checked too but are full-precision exports (radarwith 12-13 decimals against+96=2), so they cannot test a display setting either way. Vendor source S4 (GXDB.py, re-fetched this session) documentsget_chan_width("a channel's display width") andget_chan_decimal("number of digits displayed to the right of the decimal point") -- names for exactly these two fields. - Line record
+0is the vendor'sDB_LINE_TYPE.Llines read0(NORMAL),Tlines2(TIE), with the 9 "lowestTline" exceptions from Session 9. Those exceptions dissolved in step 4. - Line record
+28flags the line after a repeat line. All 8 instances (SAMAGEM_CDI,MLGRAV) sit on the line whose number directly follows a dotted.1repeat:L1004.1->L1005,L1263.1->L1264,L1267.1->L1268,L1269.1->L1270,L1291.1->L1292,T127.1->T128,L2180.1->L2190,T4060.1->T4070. A field of one record describing its neighbour suggested the record boundary itself was wrong. - Hypothesis: every symbol record begins with its name. Then our line
+0..+31would be the tail (+96..+127) of the previous line's true record, and our channel/user+0..+7the tail (+120..+127) of the previous true record. Tested against the geometry first, with no free parameters: with true line table startL = c0 + 8 - lines_max x 128and true blob-symbol startL - blobs_max x 128, on 22 of 22 files the directory ends exactly where the blob-symbol table begins (280 + word44 x 6), the line table ends exactly where the channel table begins (c0 + 8), and the user table ends exactly on header word - The "24 bytes (unexplained)" after the line table, the "8 bytes" before word 104, and the "first 32 bytes of the blob-symbol table coincide with the last five directory slots" overlap (spec section 2.1) all disappear: directory, blob symbols, lines, channels and users are one contiguous run of 128-byte records.
- Field-level predictions, all held. Reading lines at the true
boundary: type at true
+96matches the name prefix on 5,003 of 5,003 lines (L->0 4,769,T->2 234; the 9 exceptions were each the firstTline reading the lastLline's type); true+124equals the dotted suffix on 5,003 of 5,003 (9 lines.1->1, 4,994 undotted->0). The channel true tail+120..+127is zero on 602 of 602 real channels. Blob-symbol records now read a clean name at true+0. - Vendor corroboration (S3,
gxapi/__init__.py, re-fetched).DB_LINE_LABEL_FORMAT_LINE/VERSION/TYPE/FLIGHT/DATEname exactly the five per-line values the true record holds (number true+92, version+124, type+96, flight+100, date+84);GXDB.line_datereturns a float (the decimal-year float64).DB_CATEGORY_BLOB_NORMAL = 0,DB_CATEGORY_USER_NORMAL = 0. - The blob-symbol table names every administrative object. Named
records:
Line Selection,Display List,Database Extension Objects,__dbreg(22 of 22 files each),?|IPJ_<X>:<Y>for each projection object (naming its coordinate-channel pair, e.g.?|IPJ_X_Rx:Y_Rx), and__<n>for per-symbol REG objects. True+76is0(DB_CATEGORY_BLOB_NORMAL) on exactly the 1,267 records that have an administrative blob atdata_slots + slot; all 1,050 name-bearing records with bit0x10000set have none (leftover names in freed slots). True+84equals the object blob's own+24size field on 1,265 of those 1,267. __<n>is a global symbol handle, and it correctsfind_channel_settings. Header words 72/76 are the running totals blobs / blobs+lines (section 6.1b), so handles[word76, word80)are channels (handle - word76= channel slot) and[word72, word76)are lines.find_channel_settingscurrently attaches an object to channelslot % chans_max(the decomposition of itsblob_index). Oracle: a REGLABELequal to a real channel name. The handle mapping names that channel 218 times, the modulo mapping 0 times; 18 match neither -- 16 are derived channels whose label was copied from their source (MGA_EastlabelledEASTING,__XlabelledX), 2 are oneDB_Mag_833.gdbobject at handle2620 == word72. Against the.dfnUNIT=for AG106386: handle 5 of 7, modulo 0 of 7. Examples of the modulo mapping's errors:LABEL "Line number"->Fid,"WGS84 Longitude"->AltEll. Of 1,732__<n>REG objects, 988 have channel handles and 744 line handles; 19 of the line-handle ones hold flat keys (UNITS,LABEL,FORMULA,_PJ_*,DB_CHAN_Y/Z), which a line-level setting would not obviously have -- [UNKNOWN] whether these are real line objects or stale handles from an earlier table size.- Header words 8..20, 68 and 116, rechecked. Words 8..20 read
(268566528, 264, 0, 0)and word 68 reads0on 22 of 22 files -- unchanged, no new angle. Word 116 (26966 onSAMAGEM_CDI, 26 onMagnetic_Data, 0 elsewhere) was compared against the pages of chain blobs no directory slot references (34,852 / 172; and non-zero on every file where word 116 is zero) -- refuted as a lost-page count.
What it changes. Session 10 item 3's "channel identity does not
determine content kind" used the modulo channel mapping and is withdrawn
(untested under the handle mapping). The Session 9 line-record +0/+4/
+8..+27/+28 and user/channel +0..+7 fields are re-read as the tail of
the preceding record. find_channel_settings attaches settings to the wrong
channels on every corpus file; not fixed this session (reported to the
user first). Docs updated: docs/spec.md sections 2.1, 3.1-3.3, 9, 11;
docs/provenance/notes.md new section 6.2d, notes in 6.3b and 6.8c.
Follow-up: the fix ("fix it!"). Added pygdb.gdb_reader.read_blob_symbols
(live blob-symbol names, category 0 only) and switched
find_channel_settings to attribute each REG object by its __<n> handle,
skipping line handles and other names. Its max_real_line_slot parameter
went away: administrative blobs are now exactly blob_index >= data_slots.
The test builder build_real_layout_gdb_bytes gained blob-symbol records,
header words 64/72/76/80 and the true 32-byte line-table position; all
real-layout tests still pass with that position. Corpus check after the
fix: of 121 channels whose returned LABEL is a real channel's name, 111
name themselves and 10 are derived channels labelled after their source.
The old integration assertions (ch_11 FORMULA, base LABEL on
Magnetic_Data.gdb) came from the wrong mapping and were replaced. The
ch_11 object carries a line handle (__1541); the base label
belongs to time. One new unknown: on AG106386, _PJ_* keys
(_PJ_IPJ = IPJ_Easting:Northing) also land on Fiducial, GPS_Height
and Ground_Speed through live channel handles -- possibly orphaned
objects whose handle was reused, not yet distinguishable (notes section
6.2d).
Follow-up: stale spec text and the suspect LABEL/UNITS tests. Fixed
two stale spots in docs/spec.md. The section 11 65636 entry now
reflects the freed-slot reading. Section 9 no longer lists "which channel
a channel_slot concerns" as open. Re-ran the four LABEL/UNITS
population hypotheses from section 6.8c, which had used the modulo
channel attribution, corpus-wide under the handle mapping
(labelunits.py; 761 LABEL and 646 UNITS slots):
- dtype, array-ness and display format are still mixed, so still refuted;
- data presence turns out untestable, since every attributed channel has
data, and its old quadrants were mapping artefacts;
- a new "live directory copy" hypothesis is mixed too.
The new finding is that population is strongly per-file. LABEL is filled
on every object in six files and on none in four, which suggests the
writing tool. Recorded in notes section 6.8c.
Follow-up: the flat-vs-nested selector, re-tested with the correct
owner ("Can you test now what selects flat vs nested content").
Classified every REG object's content against the owner its blob
symbol names (selector.py..selector5.py).
- Every nested object is a
MAKERrecord. 24 of 24, all on channel handles, naming the creating GX (newchan.gx20,linechan.gx4). Every live__<n>handle names exactly one object (1,109 of 1,109). - Many "numeric" objects were really flat objects with slot 0's first
4 bytes zeroed. 107 of 110 tails equal a known key minus its first
4 characters (
S,L,ULA,DATUM_TRANSFORM,HAN_X). This explains the Session 10 "dirty slot 0" byte as the key's 5th character. - Preamble
+124is the selector. It is non-zero on exactly the 860 objects whose+128starts a key. There it equals the count of consecutive key slots from slot 0 (860 of 860). It is zero on all 898 others, where+128is1+ff 00 f0 0ffor the nestedMAKERform (24) and0otherwise. - Line-handle objects are never clean flat or nested. They are only empty, binary, or zeroed-key objects.
Reading: an object rewritten with zero entries keeps its old bytes past
the header, and 8 objects also carry key slots past their count.
find_channel_settings currently returns keys from 7 channel-owned
+124 = 0 objects and from those surplus slots -- not changed yet,
reported to the user. Recorded in notes section 6.8c and spec section 9.
Follow-up: the reader honours +124 ("yes").
_decode_reg_flat_keyvalues now reads only the first +124 slots; a
count of 0 means no flat entries. Test fixtures write the count, and two
new tests cover the real leftover shapes. Corpus effect on
channel_settings:
AG1063867 → 4 channels,DB_EM_2934 → 2,SAMAGEM_CDI10 → 5 (its 6 conflicting-value warnings gone); label oracle unchanged at 111/10.- Every dropped
_PJ_*entry had sat on a non-coordinate channel. After the change, projection keys appear only on real coordinate pairs. That resolves the Session 11 "orphaned handle" caveat as leftover slots, and independently supports reading the dropped bytes as stale.
Follow-up: working through the remaining open questions ("keep trying
to tackle the remaining questions in the order you see fit"). Scripts
lineobj*.py, numeric.py, blobhdr*.py, dirflag*.py, freelist*.py,
admintags.py.
- Line-handle REG objects. All 744 declare zero entries, and each handle is a real line slot (line 0 in most files; 417/631 and 283/631 lines in the two USGS files). Tested the idea that these are channel objects shifted by a table resize -- no constant offset fits, so it is refuted. Reading: always-empty per-line registries whose surviving bytes are leftovers. Notes section 6.2d.
- Binary
+124 = 0content: numeric array or leftover? 48 of 117 probe runs recur elsewhere in the same file, mostly in other registry objects -- inconclusive, stays open. - Blob-header
+20is a blob class. 100 admin, 200 plain data, 202 compressed data (10,429/10,429 vs 106,681/106,681).+16is a write timestamp set on whole imported channels (Magnetic_Data: 9 channels x 631 lines on 2020-03-25, the line date), sentinel otherwise. Notes section 6.1e. - The "cache" directory slots are a free list, and header word 116 is
the lost-page count. No free-list entry points at a referenced blob
(0 of 117,450). In 20 files every orphaned blob is listed and word 116
is 0. In the two files with a full list, word 116 equals the unlisted
orphans' pages exactly (26,966 and 26), and the supplied file agrees.
This corrects this session's first-round "refutation" of word 116,
which counted listed orphans too. It also explains the unlisted data
blobs as freed. The
0x4flag bit is still unexplained. Notes section 6.1d. - Every administrative blob is identified by its symbol name. The
+44tag follows the name exactly (REG, IPJ, EXT, META,ff ff ff fffor Line Selection, a varying int for Display List, plain text forOE.DB_ACTIVITY_LOG/OE32.View). The unidentified GSQ "third tag" isOE.DB_ACTIVITY_LOG. Display List holds channel name + handle pairs. Notes section 6.2d. - Channel-record census: leftover records read as channels (reader
fix). 16 records in
DB_AGG_1213,DB_Mag_1213andDB_Mag_1212passedlooks_sanealthough they have array width 0, dtype 0 and no blobs. Their names are second copies of real channel names, fill patterns, and projection-catalog names.GDB.channel_nameslisted them.looks_sanenow also requires array width >= 1; there is a synthetic and a corpus regression test, and the three files now list 16 unique channels each. Genuine channels all read+108 = 1.0and+116 = 5(535 of 535, meaning unknown). The freed-slot bit reading does not transfer to channels: 53 genuine channels share the int320x10000at+84. OE.DB_ACTIVITY_LOGis a creation record. It appears only in the four Melinda Downs files. It holds the source path, aCreated:timestamp (2008-05-08), andLines:/Channels:equal to the file's ownlines_max/chans_max. The blob header overwrites the start of the text. It does not cover the 1991 files whose declared Speed mode was never used, so that question stays open.- Compressed-chunk
reservedtracks the subtype. It is 0 on all 7,015 LZRW1 chunk headers and 1 on all 3,529 zlib ones,.gdband.grdalike. The five largest.gdbfiles were skipped for memory. The two stray values are the magic occurring inside data. Line Selectionis one byte per line slot, and blob+24looks like a payload length from+28. Every liveLine Selectionobject holds exactlylines_maxrounded up to 4 bytes offffrom+28, and+24equals that count. Tested on registries: empty ones end at+132(874/874), and flat ones end 4 bytes past their last slot (678/860). The rest are unexplained, so this is recorded as [LIKELY], partial.-
The registry grammar ("keep going on anything that doesn't require new files"). The 182 flat registries that missed the
+24rule each have an int321after their slots, then a nestedMAKERobject. The payload ends on that object's0x1Abyte. The 678 hits have a0there. So every registry is:nentries (+124), an int32 countm, thenmnested objects. Every frame length (+32,+64, the nested frame's) counts from 28 bytes after itself and ends where+24says.Over 1,820
REGobjects (corpus plus supplied file): frames 1,820/1,820;m = 01,600/1,601;m = 1219/219. The "nested", "empty", "numeric array" and "dirty slot 0" shapes are all this grammar withn = 0, and the last two lie outside the declared payload, so they are leftovers. That also settles the "numeric-array-or-leftover" question.The one exception is a supplied-file registry with an
ASSOCIATED.DB_TABLE$$$entry whose table follows as an extra"REG "block. Reader comments and docstrings updated to the grammar (no behaviour change). 11. Frame markers and the IPJ member chain.ff 00 f0 0fandff 00 e1 1eare the only self-complementary markers in any administrative object. Each is followed by a length that counts from 28 bytes after itself, andf00fframes enclosee11eframes. So they read as object / member, which explains the preamble "constants".IPJ objects hold a chain of 2-6 members that lands exactly on the payload end (63/63). Member 0 is the known projection record; member 1 is an unused
-1e32array; members 2-3 are zero. East_Isa's extra member 4 carriesEPSG+ 28354, the EPSG code of its own name "GDA94 / MGA zone 54" (one instance, [LIKELY] authority code). 12. IPJ projection parameters are an 8-slot vector at+588. Comparing member 0's float64 run with the_PJ_PROJECTIONtext: the text's five Transverse Mercator parameters fill slots 0, 1, 4, 5, 6, in text order.+588(slot 0) is the latitude of origin, always0here -- a new field.+604/+612/+644are the unused slots 2, 3 and 7, and the vector ends exactly at member 0's end (+652). 13. IPJ name field and type word. The name at+104sits in a 64-byte field (all 63 fit), and the "pointer-shaped"+136..+176bytes are uninitialized memory after its NUL. The old "+112grid-system tag" was name text and is withdrawn.+168is11on all 48 projected objects and1on all 15 datum-only ones. It does not matchIPJ_TYPE_*, and its meaning is unknown. 14. User-record path: layout and truncation rule. The path is UTF-16LE from+8, overwritten at its start by the ASCII user name, and at most 32 characters long (to+71). It holds on 11 of 22 files, each ending in its own file name. Shorter strings are NUL-terminated (7); longer ones are cut at 32 characters with no terminator (4). This corrects the old "32-byte field at+40, 13 of 22" (the USGS fragments are not paths) and the "32-byte name field" inference. 15. The directory0x4bit flips on every rewrite. Pairing each live entry with the freed copy of the same blob index: opposite bits in 696 of 696 pairs (404 live0xC/freed0x8, 292 live0x8/freed0xC), and objects with two freed copies have one of each. Never rewritten entries are0x8. Magnetic_Data's 26 unpaired live0xCentries match its 26 lost pages. So this reads as a write-generation parity bit ([LIKELY]); the purpose is still open. 16. Thef0f0f0f0header variant is at bytes 4-7, not 8-11. A direct header dump of both files shows the variant in header word 4; Session 2's note and the spec had the offset wrong. The Elaine header differs from its same-delivery siblingDB_Mag_MountGordon_1003only in that word and the page count, so nothing else in the header explains it. The meaning remains unknown. Corrected in spec section 2, notes section 6.1, andcheck_magic's docstring.
Session 12 -- new public samples; the first non-Transverse-Mercator projection (2026-09-29)¶
Prompted by: "let's try to find more gdb files online, I don't think we need reference files anymore though do we?" Paired reference exports are no longer needed for the value decoding. What the open questions need is structural variety: other projections, several users, other line types, channels with non-default settings.
- Search. Web search for public Geosoft databases turned up USGS OFR
2006-1204 (Afghanistan gravity), which ships
afgrav.gdbdirectly (435,200 bytes,pubs.usgs.gov/of/2006/1204/Gravity/; thepubsdata.usgs.govhost named on the page does not resolve). Its readme states a Transverse Mercator projection with "Base latitude = 34 degrees N". Downloaded with the readme intosamples/usgs_afghanistan_2006/(gitignored). - IPJ
+588confirmed as latitude of origin. The Transverse Mercator object reads 34.0 there, as the readme predicts. - First non-TM projection: Lambert Conic Conformal (2SP), method
code
+168 = 3. Its eight parameter slots follow the registry text"Lambert Conic Conformal (2SP)",30,38,0,66,0,0in order (slots 0-3, 5, 6), where Transverse Mercator uses 0, 1, 4, 5, 6. So+168is a projection-method code (1 geographic, 11 TM, 3 LCC), and the slot meanings depend on it. - Reader defect:
find_projection_parametersreads TM slot positions for every object, and reportscentral_meridian=38(a standard parallel) for the Lambert system. Reported, not fixed yet. - Line type 6 (
RANDOM) on lineD0, and empty user slots read0x10000, like the other symbol tables.
Docs updated: spec sections 3.2 and 8; notes section 1 (S16), section 5,
and 6.7b.
6. More files from pubs.usgs.gov directory listings. OFR 2004-1096
(Long Valley) has a browsable downloads/geosoft_gdb/ directory with
long_valley_ed.gdb (4.6 MB). OFR 2011-1270's report/ directory
holds six small databases of Afghanistan ground data digitized from
Soviet maps. All were downloaded into samples/ (gitignored); the
corpus is now 30 public files plus the supplied one.
7. Re-check with a survey script. survey.py re-tests every decoded
rule on a file. All eight new files pass all of them.
long_valley_ed.gdb gives a third non-zero header word 116 (1,148),
equal to the unlisted orphans' pages, as predicted from the free list.
8. Group lines carry a group class. All 110 lines of the six OFR
2011-1270 files are category 200 (group), with freeform names. Their
date/number positions hold the NUL-terminated class name DB_Table.
Vendor S4 set_group_class documents exactly this ("the Class name
for a group line"; group lines of one class share associated
channels). The associated-channel list is __dbreg's
ASSOCIATED.DB_TABLE key, a comma-separated channel list. It is
present only in these six files, and its $$$ twin is an ordinary
empty entry.
9. Freed user slots carry 0x10000 in every new file, including
slots holding leftover names (SPF_1, SPF_250).
10. Two reader bugs from the new files, fixed. Kalay_nk.gdb has
the corpus's first live data directory entry with the rewrite bit
(0xC). It points at the earlier of two copies, and the later copy
is on the free list. The reader accepted only 0x8, so it warned and
fell back to the later (freed) copy. Here the two decode to identical
values (they differ only in padding after one record's NUL), but the
logic was wrong. BlobDirectory.resolve now accepts any entry with
the live bit and masks the page with 30 bits.
The same channel (`string[255]`) exposed a native-backend crash on a
string blob whose records are all empty (`<U0` dtype). It now returns
empty strings like the Python path. Both fixes have regression tests.
The integration tests were generalized: unlisted blobs must be freed
(on the free list or counted by word 116), word 116 is checked on
every file, and "has REG/IPJ" accepts a `__dbreg` registry
(`Kalay_nk`/`Kalay_pk` have no coordinate system).
- Projection reader fixed ("fix the projection reader, then keep
searching").
find_projection_parametersnow reads the method code at+168and maps parameter slots per method: Transverse Mercator 0/¼/⅚, Lambert Conic Conformal (2SP) 0/½/⅗/6. It no longer names fields for unknown methods.ProjectionParametersgainedmethod_code,latitude_of_origin,standard_parallel_1,standard_parallel_2and the rawparameterstuple. Onafgrav.gdbthe Lambert system now reads parallels 30/38, origin 0, central meridian 66 (wascentral_meridian=38). Synthetic tests cover TM with non-zero origin, Lambert and an unknown method, and an integration test uses the real file. - Alaska DGGS GPR 2015-4 (Fortymile). The survey's Geosoft-database
zip (115 MB) is a direct download from
dggs.alaska.gov/webpubs/data/.fortymile_linedata.gdb(173 MB,DB_COMP_SPEED, NAD27 / UTM zone 7N) passes every rule. It contributes the corpus's first non-WGS84 ellipsoid: Clarke 1866, semi-major axis 6,378,206.4 m, eccentricity 0.08227185, read correctly at+308/+316. - Session resumed after a usage-limit reset ("keep on chugging"). Finished the BAS doc edits. The shell's safety check was briefly timing out, so they went in through the editor.
- More files.
- OFR 2006-1204's
German_mag/GDR_clmag.gdb(41 MB) adds the International 1924 ellipsoid (Herat North datum). - OFR 99-527 (Wisconsin) has only ASCII/binary line files.
- NRCan's portal fails TLS certificate verification. Not bypassed; skipped.
- Alberta's IAM-013 links to grids only.
- OpenEI GDR 1682 (BRIDGE, CC-BY 4.0) packages
.gdbfiles inside multi-GB zips. A range-request zip reader (remotezip.py) listed the members and extracted the 12 Bell Flat databases (2024) without downloading the whole archive.
- OFR 2006-1204's
- Findings from them.
- Every rule holds.
Line Selectionroundslines_maxup to a multiple of 8. It has no tag at+44: itsffbytes run through there, and a short payload leaves leftovers there.- The
f0f0f0f0header variant is on all 11 BRIDGE ground-gravity databases, 13 of 45 files in all. It follows no tested file feature, so it stays unknown.
Session 13 -- decoding the remaining administrative objects (2026-09-29)¶
Prompted by: "Let's try to push on decodable contents from the files we
have." Scripts dumpobj.py, vv.py, maker.py.
- Display List. It is a bare VV object:
00 1a cc ff+VV, int32 0, int32 -82, int32 count, then count x 82-byte records ofchannel name\0handle\0. It ends exactly at the payload end. - The VV is a fixed-width string vector. The element type is minus
the width, as in channel string types. It is -256 on every registry
(the old "
00 ff ff ffconstant" at+120), and -82/-130 onDisplay Lists. The count is the VV length: a registry's+124.Display Listhandles identify channels, and its names are cached labels (renamed channels keep old names). - EXT is an empty list: class
EXT, a member with codeLMSL, no content, identical everywhere. Registry-like bytes after it are leftovers outside the payload. - Preamble generalized.
+44is the object's class name inside the object frame, and+96the member's code inside the member frame. - MAKER decodes on 305 of 305: tool name, a 2-byte zero field,
label, then
TOOL.KEY="value"parameters. The 2-byte field was found when 4-byte alignment fit one tool but not another. All 32newchan.gxrecords carryDISPWIDTH/DISPDIG/ARRAYSIZE/NAMEequal to the channel record's+94/+96/+118/name, which upgrades channel display width/decimals to [CONFIRMED]. __dbmeta. Container decoded (classMETA, memberATEM, an 11-int header ending with the zlib and raw lengths, then zlib). The content is a per-file typed metadata tree of vendor types and a snapshot of the database's channel metadata. The node-record grammar is left undecoded (3 of 47 files).
Session 14 -- checking whether channel +116 tracks dtype (2026-09-29)¶
Prompted by a coordinator question before treating channel +116
(always 5 on 535 of 535 corpus channels, per Session 11/§6.2d) as a
genuinely independent constant: had it only been checked on
GS_DOUBLE channels, where 5 also happens to be the real dtype
code? If +116 tracks +84 (the channel's own dtype code) it is a
second copy of the type, not an unexplained field.
Checked directly with pygdb.gdb_reader.read_channels() against every
.gdb under samples/ (48 files, all parse). Grouped every channel's
+116 value by its own dtype_code:
dtype_code=-255..−9 (9 distinct string widths) -> +116 always 5
dtype_code=1 (GS_USHORT, incl. Radiometric_Data.gdb's 512-wide
ISPD/ISPU spectrum channels) -> +116 always 5
dtype_code=2 (GS_SHORT) -> +116 always 5
dtype_code=3 (GS_LONG) -> +116 always 5
dtype_code=4 (GS_FLOAT) -> +116 always 5
dtype_code=5 (GS_DOUBLE) -> +116 always 5
Result: refuted. +116 reads exactly 5 regardless of the
channel's real dtype -- including on GS_USHORT array channels whose
own dtype code is 1, not 5. It is not an echo or second copy of
+84; whatever it is, it is genuinely independent of the channel's
type. docs/spec.md §3.1's "[UNKNOWN]" label for +116 stands as
before; this session only strengthens the "independent constant"
framing with an explicit non-GS_DOUBLE cross-check, since the
existing "535 of 535" figure did not previously call out that it spans
every real dtype in the corpus, not just GS_DOUBLE.
Session 14 -- the coordinate-system definition, from vendor documentation (2026-09-29)¶
Prompted by: "On the coordinate systems, I see there's a Coordinate_system class in the gxpy source code (again don't run anything). and seequent's public GX developer wiki has a dedicated page on them. Does those help you?"
- Sources read, nothing run.
gxpy/coordinate_system.py(S24) is a wrapper over the compiledGXIPJAPI. It defines a coordinate system by five GXF strings and points to the GXF specification for them. The wiki page (S24) gives the same strings. GXF Revision 3 (S23, a Geosoft document hosted by USGS) defines them and, in Table 1, lists every projection method's parameters in order. - Slot names confirmed. Table 1's parameter order, with unused EPSG parameters kept as unset slots, names every binary slot for the three decoded methods. The Lambert and Polar Stereographic names go from [LIKELY] to [CONFIRMED].
- The rest of IPJ member 0 found, using the registry's GXF text as a key:
- prime meridian
+324; - the 7 Bursa-Wolf transform parameters
+396..+451(rotations in radians, scale as1 + ppm/10⁶, exact against three datums); - units name
+452and factor+516; - projection name
+524.
Every field's offset now tiles +104..+652 without gaps.
4. Registry text vs. binary slots, corpus-wide (pjtext.py). For
every projected IPJ object, _PJ_PROJECTION values were collected
from the registries' flat entries under a matching _PJ_NAME (the
value is quoted). 43 of 50 projected objects have text; 7 have none
(three in AG106386, DB_Mag_1212, DB_Mag_1213, both systems in
Brunt_mag_2017). On all 43, the text holds exactly as many values
as the binary has set slots, equal in order (41 Transverse Mercator,
1 Lambert (2SP), 1 Polar Stereographic). No geographic object has
_PJ_PROJECTION text. This is what lets the text name the slots of a
method whose code has no known layout. The reader now does so, keeping
the binary's values.
5. The "no transform" shape (pjcheck.py). All 25 datum transforms
with registry text convert back to it (arc-seconds, ppm) within
1e-6. The 5 objects whose transform name field is empty (the three
GDA2020 systems in AG106386, NAD83(2011) in
HawthorneGrav_wMasks_MF20240116, NAD83 in long_valley_ed)
hold dX = rDUMMY (−1e32), dY..Rz = 0 and scale = 1. A WGS 84
datum instead holds all zeros and scale 1. The empty name field can
hold leftover bytes after its leading NUL.
Session 15 -- baselines for the unseen-feature notices (2026-09-29)¶
Prompted by: "are we at the stage where we could notify the user if their
file has something we haven't seen yet that would help us narrow down
those last few pieces?" The notices themselves are a reader feature
(pygdb.unseen). Checking each "always seen" fact before building on it
was format work. Scripts: baselines.py, notices.py. Recorded in
notes.md section 6.2e.
- Header words 8-20 and 68. Tallied as int32 on all 49 files (48
public plus the supplied file): one combination only, word 8
0x10020000, word 12264, words 16/20/68 zero. - Channel
+108/+116. 1,247 genuine channels read 1.0 and 5. Every one. - User table. Every file has exactly one live user record, category
0x20000,+124= −1. - Line records, read at true offsets. Of 5,575 live lines, the types
are 0, 2 or 6, and the 20-byte block at true
+104..+123is the same bytes on all of them. The spec's earlier "all-zero on 18 of 5,003" came from the previous-line view. It does not reproduce at true offsets, and the 18 were not traced. - IPJ members 1-3. My first member walk was wrong (next member at
offset + 8 + 28 + length, which never lined up). Printing the frame positions (652,776,872after a 560-byte member 0 at60) gaveoffset + 32 + length, which ends exactly at the payload end on all 107 objects. From each member's 16-byte index onward, the bytes are identical everywhere: - member 1: 28 zero bytes and eight
rDUMMY; - members 2 and 3: 64 zero bytes.
The first draft of the check assumed 40 zero bytes and six dummies, then 44 and six; both fired on every file until the bytes were dumped and counted. The 8 bytes between the length field and the index vary between objects. 6. Nothing in the corpus trips any check: 0 notices over all 49 files with channels, lines, projections and makers read. This is now an integration test.
Session 16 -- four new samples surface a real bug in the unseen-feature checks (2026-09-30)¶
Prompted by: "Can you search for more gdb files online, for us to test against?" and "after downloading them, first run them through our new helper function to see if they trigger any of the warninigs."
- New source: Geoscience BC's Golden Triangle compilation (project
2018-056, report GBCR2021-15,
cdn.geosciencebc.com). Ten legacy contractor surveys (2013-2020), each area's zip a few hundred MB to a few GB,accept-ranges: bytessupported. Usedremotezip.py(as for the OpenEI BRIDGE zips, Session 12) to list and extract single small.gdbmembers without downloading a whole zip. Addedsamples/geosciencebc_golden_triangle/:Area_I1_airborne_MAG_2020(39 MB),Area_H1_MAG_2020(50 MB),Area_G1_MAG_VTEM_2013(81 MB),Area_F1_MAG_2017(104 MB). - Dead ends, not pursued further: the Ontario GeologyOntario hub
domain no longer resolves at all (DNS failure, was merely
JS-blocked before); GeosoftInc/gxpy's own
test_free_gdb.pyfixture is explicitly out of scope (Session 1.12: engine-generated binary output, even in a public repo, is not a clean-room source); modern USGS EarthMRI ScienceBase items now bundle everything (report, grids, database) into one multi-GB zip, no smaller Geosoft-only sub-item; South Africa's Council for Geoscience and Brazil's CPRM have no direct.gdbdownloads found. Area_F1_MAG_2017.gdbimmediately surfaced a real bug inpygdb.unseen, per the new "run the report on every new file" instruction: aprojection method code 0notice, with an empty coordinate-system name ('?') and every parameter slot unset.- Not a new format discovery.
_live_admin_objects(whichunseen.check_admin_objectsscans) surfaces every liveIPJ-tagged, gated object -- including a kindfind_projection_parametersalready silently skips: an object with no decodable embedded name. Checking the corpus (_live_admin_objects+ the gate check, 46 of 49 files) found 133 such gated objects. All 118 with a decodable name have one of the four known method codes (68 TM, 43 geographic, 5 PS, 2 LCC); all 15 without one read method code 0 -- a clean split, not a coincidence. Their blob-symbol names are?|IPJ_<channel>:<channel>(e.g.?|IPJ_x:y,?|IPJ_Longitude:Latitude), one per (channel, channel) pairing -- real, common objects (46 of 49 files, so present in the corpus all along, just never scanned by symbol name before), evidently a per-channel-pair placeholder distinct from the named working-coordinate-system object, not a genuinely novel projection method. - Fixed:
unseen._find_ipjnow skips an object with no decodable name, exactly likefind_projection_parameters. Also tightened the gateunseen.check_admin_objectsapplies before trusting anIPJ-tagged object's fixed offsets at all, to require the" JPI"marker at+96(matchingfind_projection_parameters's own gate) -- not the cause of this particular false notice (the real objects do carry the marker), but a latent gap the same investigation surfaced. Two regression tests added (tests/test_unseen.py): a gated object with no name (must not fire), and a same-shaped, ungated object (must not fire either). Full suite re-run clean on all 285 tests, including the corpus-wide "no notice on any real file" test now covering the four new samples too. - All four new files otherwise pass every existing rule.
Area_G1_MAG_VTEM_2013.gdbhas one already-known, already-handledGDBParseWarning(a channel's creation record has two differing stale copies) -- not a new finding.