Skip to content

feat: read gzip-compressed files as the file inside them (#68) - #80

Merged
vaceslav merged 15 commits into
mainfrom
feat/gzip
Oct 3, 2026
Merged

vaceslav merged 15 commits into
mainfrom
feat/gzip

Conversation

@vaceslav

@vaceslav vaceslav commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

What and why

Closes #68. Part 1 of #76 (more archive formats). TabularFile.Open now reads a gzip-compressed file — data.csv.gz, report.xlsx.gz, table.ods.gz — as the file inside it, and refuses a cut-off or damaged one instead of reading it as a shorter file.

Spec: docs/superpowers/specs/2026-10-03-gzip-design.md. Plan: docs/superpowers/plans/2026-10-03-gzip.md.

  • Detected by bytes (1F 8B 08), opened as GzipCursor. FileProfile.Format is Gzip; each sheet keeps the inner file's Format. A csv is decompressed as it is read, never unpacked; an xlsx/ods is decompressed into memory once, under ArchiveCursorOptions.MaxEmbeddedWorkbookBytes. What a file expands to is bounded by MaxUncompressedBytes, counted while decompressing.
  • SheetInfo.Source is the file name the gzip header stores (FNAME), else null — never the name passed to Open, because a mapping plan records Source and the name a caller passes may differ between analysis and import. The csv sheet is named after FNAME, else the file name without .gz.
  • Our own gzip framing over the BCL's DeflateStream. Measured on .NET 8.0.11 and 10.0.9: GZipStream reads a gzip file cut off anywhere as a prefix of its content and reports success, so a cut-off upload would be imported as a smaller, valid file. GzipHeader, GzipStreamReader and Crc32 (managed slicing-by-8, ARM64 instruction where present) check every member's CRC-32 and size; cut off → format.truncated, damaged → format.corrupt. Several members (cat a.gz b.gz, bgzip) read as one file; trailing bytes that do not start another member are ignored, as the gzip tool does.
  • Refused: a zip archive or another gzip file inside a gzip file. A .gz inside a zip archive is skipped as the new SkippedEntryReason.Compressed (was Binary).
  • EmbeddedWorkbook (internal) is now shared by ArchiveCursor and GzipCursor; ArchiveCursor's behaviour is otherwise unchanged.

Performance (5M-row csv, 572 MB; Release, one run each, machine under load):

csv csv.gz zip zip before
read 3.12 s 4.30 s 4.22 s 4.25 s
full analysis 13.1 s 14.6 s 14.3 s 14.2 s
peak RSS (read) 54 MB 55 MB 57 MB 57 MB

The .gz profile equals the plain file's (row count, all 17 columns, repair diagnostics). The final reviewer measured the decompression path at GZipStream's speed on a single member (200 MB: 703 vs 706 ms on net10) and 1.3–1.6× on bgzip-shaped files of 64 KB members.

Tests: every cut of single-member, two-member and empty files; the empty member a flushing writer produces (Python, zlib sync flush); a member whose deflate data ends in tens of thousands of empty stored blocks; wrong CRC, wrong size, broken deflate data; bounds (a gzip bomb, an inner workbook over budget); xlsx/ods/csv inside; detection near-misses; analysis → plan → import under a different name; gzip in the stream-ownership, cancellation, API-contract and null-argument suites; a gzip fuzz case (20,000 cases run locally). Independent probes by the reviewers: ~16,000 cuts and 8,000 bit flips over 400 random multi-member files, none read silently.

Known, accepted: a cut exactly at a member boundary — or one that removes only the last trailer byte of a final empty flushed member — reads as the complete shorter file; the bytes cannot tell them apart and no content is lost.

Checklist

  • A test that failed before the change and passes after it (for a fix or a new behaviour).
  • dotnet build and dotnet test pass on net8.0 and net10.0 with zero warnings. (2288 tests, rebased on main)
  • If the read path changed: measured, and the numbers are in the description. (existing read paths unchanged; the new gzip path and zip before/after measured above)
  • If an error code was added or changed: ErrorCodes and the guide's error-code table agree. (no new codes; existing format.truncated / format.corrupt / format.unsupported)
  • Public API changes are described, and breaking ones are called out. (additive: GzipCursor, TabularFormat.Gzip, SkippedEntryReason.Compressed; a consumer's exhaustive switch over either enum needs the new member)
  • No third-party package in src/.

The narrow window misses the end of the deflate data when it lies in a later read of the final decompressor call; search everything that call read, in chunks, before refusing. Also docs: Source, bounds, README period.
@vaceslav
vaceslav merged commit 78e1223 into main Oct 3, 2026
9 checks passed
@vaceslav
vaceslav deleted the feat/gzip branch October 3, 2026 19:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Read gzip-compressed files (.csv.gz, .xlsx.gz)

1 participant