mirror of
https://github.com/xroche/httrack.git
synced 2026-07-23 09:09:05 +03:00
* Decode the br and zstd content codings, and bound every decode httrack advertised gzip and deflate only, so it took the oldest coding on offer while every browser negotiates brotli (46% of CDN traffic) or zstd. Worse, the response side had no codec identity at all: Content-Encoding was collapsed into a boolean and hts_zunpack guessed the framing from the body bytes, so a br or zstd body sent unsolicited (mis-keyed Vary caches do this) fell through the identity path and was saved as the page, coded bytes and all. Content-Encoding now maps to a codec (htscodec.c), and the decode dispatches on it: brotli and zstd get streaming decoders, a known coding we cannot undo fails the fetch instead of saving garbage, and an unrecognized token still means identity, because servers do put charsets in that header. br and zstd are advertised over TLS only, as browsers do. On any decompression failure the undecoded body is now dropped rather than left in memory for the writer to commit as the page. That closes the coded-bytes-as- page hole for the new unsupported-coding path, and with it a latent leak on the gzip path, where a body that failed to inflate (a truncated stream) was written to disk verbatim. Every decode is now bounded: 4096x the coded body, floor 1 MiB, ceiling INT_MAX. Nothing capped the decoded size before, which was survivable while deflate was the only codec (it cannot pass 1032x) but not with brotli and zstd, which reach a million to one. The zlib accumulator also moves to LLint, so a body over 2 GiB fails instead of overflowing an int. libbrotlidec and libzstd are optional at configure time and required on Windows, where libhttrack.dll links them from vcpkg and CI runs the self-tests to prove the decoders are in. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com> * Fix copyright year on the new codec files (2026, not 1998) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com> * tests: probe the direct-to-disk decode-failure discard, and assert absence The wire test only exercised the in-memory discard path (bad.html is HTML). Add a non-HTML body under an unsupported coding (bin.dat) so the is_write branch is covered too, and assert both are absent from the mirror rather than grepping a file that the fix now removes. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com> * Factor the budget-enforced write shared by the brotli and zstd decoders The two streaming decoders duplicated the decoded-size budget check, the security-critical part of each loop. Pull it into one codec_sink helper so the bomb ceiling lives in a single place. Behavior-preserving; the codec self-test (exact decoded bytes, truncation and bomb rejection) proves equivalence. Also compile the streaming-decode body only when a streaming codec is built, which drops an unused-variable warning in the --without-brotli --without-zstd build. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com> --------- Signed-off-by: Xavier Roche <roche@httrack.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>