Compare commits

..

18 Commits

Author SHA1 Message Date
Xavier Roche
3563d4fbe0 Merge origin/master (#763 landed)
Signed-off-by: Xavier Roche <roche@httrack.com>

# Conflicts:
#	src/htsback.c
#	tests/96_local-refetch-keep.test
2026-07-27 09:48:55 +02:00
Xavier Roche
99bd9bb674 A failed re-fetch overwrote the mirrored file with the aborted read's debris (#763)
* A failed re-fetch overwrote the mirrored file with the aborted read's debris

A transfer that dies before a complete response has no body, yet the save path
still consulted r.adr, which at that point holds whatever the aborted header
read left behind: raw status-line bytes, or an empty buffer that truncated the
file to zero on macOS. Require a successful transfer, as the empty-body half of
the condition already did.

Closes #748

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: assert the failed re-fetch keeps its bytes on every platform

Test 93 filtered reset.bin out of its bucket lists because a connection killed
before the status line surfaced differently per platform. It no longer does, so
assert the resource like any other: unchanged, with the bytes pass 1 mirrored.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: make the re-fetch assertions catch a crawl that never re-fetched

The new test passed with no second pass at all, so a regression that stopped
re-fetching would have looked green. Assert the failure the fixture provokes,
give stay.bin a fresh pass-2 body so a fix that stopped overwriting anything
fails, and compare reset.bin by checksum rather than by length.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: assert the retry exhaustion, not the per-platform failure message

The message a cut connection produces depends on whether any bytes arrived, so
matching it would fail on a runner that sees none.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 09:44:28 +02:00
Xavier Roche
29dfd2df59 htsserver never returns from main(), it blocks on its own exit wait (#757)
htsthread_wait_n(background_threads - 1) subtracts one more than the count of
threads that must not be joined. Without --ppid that count is zero, so the wait
asks for a negative number of outstanding threads and the counter never gets
there; with --ppid it waits on the pinger, which by design never returns.

Wait for background_threads instead, which is what the sibling call inside
back_launch_cmd() already passes.

Closes #753

Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 09:43:51 +02:00
Xavier Roche
55765815f5 htsAddLink walks back from an empty codebase, one byte before the buffer (#767)
* htsAddLink walks back from an empty codebase, one byte before the buffer

Same idiom as the lienrelatif() underflow fixed in #729: the trim that walks
back to the last '/' starts at codebase + strlen(codebase) - 1, which is
codebase - 1 when the string is empty, and the loop dereferences it before
a > codebase stops it.

Unlike #729 there is no reachable empty value. codebase is copied from a
recorded link's fil, and no hts_record_link() call site can supply an empty
one: every fil is either a literal seed or an ident_url_absolute() /
ident_url_relatif() success return, and fil_simplifie() restores "/" or "./"
rather than leaving a path empty. The guard goes in anyway, and -#test=addlink
drives the walk directly: under the sanitize job's ASan+UBSan build the test
fails on the unfixed walk and passes with the guard. Its second case pins the
ordinary trim, which the guard leaves alone.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* review: add the case that actually notices the codebase trim

For an ordinary relative link ident_url_relatif() re-derives the directory
from the path it is handed, so deleting the trim outright left the two
existing cases green. A query-only link ("?x=1") copies that path whole, and
does catch it.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 09:40:11 +02:00
Xavier Roche
869b8479e9 structcheck() builds its rename target with an unbounded sprintf (#762)
* structcheck() builds its rename target with an unbounded sprintf

Both structcheck() and structcheck_utf8() move a regular file sitting where a
directory belongs, and build the "<name>.txt" target with a raw sprintf into a
2048-byte buffer. It stays in bounds only because of a strlen(path) >
HTS_URLMAXSIZE guard dozens of lines above, which nothing at the write site
mentions. Route both through sprintfbuff() and fail with ENAMETOOLONG, so the
bound is local.

Armed the probe: with that distant guard patched out, a 2045-byte path makes
the old sprintf write 2049 bytes into tmpbuf[2048] under ASan; the same build
with sprintfbuff() returns -1 and reports nothing.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* review: tighten the structcheck self-test and its comments

The path builder could end a path with a bare separator when the base
directory's length hit the wrong residue, so the test aborted on a long
$TMPDIR. The refusal case also asserted on a component structcheck never
creates; assert on the outermost one instead.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: drop the max-length rename case, macOS PATH_MAX is 1024

The path the guard admits (HTS_URLMAXSIZE) plus ".txt" is longer than macOS
accepts, so fopen() failed there. The case could not tell a fixed build from an
unfixed one anyway; what is left covers the guard and the rename on both entry
points.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 09:34:02 +02:00
Xavier Roche
45630f63b4 Merge fix/refetch-debris
Signed-off-by: Xavier Roche <roche@httrack.com>

# Conflicts:
#	tests/96_local-refetch-keep.test
2026-07-27 09:26:02 +02:00
Xavier Roche
2ffe8d5582 tests: assert the retry exhaustion, not the per-platform failure message
The message a cut connection produces depends on whether any bytes arrived, so
matching it would fail on a runner that sees none.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-27 09:25:38 +02:00
Xavier Roche
14cc7a13fc Keep the copy for every failure, not just the five retryable codes
A malformed status line, an oversized declared length or a mid-flight abort all
land on STATUSCODE_INVALID, outside the retryable set, and the purge still ate
the file. Invert the test: any failure keeps the copy except the codes that
mean the engine passed the resource over on purpose.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-27 09:23:16 +02:00
Xavier Roche
1a7effce49 --purge-old deleted a file whose re-fetch never got a response
A transfer that dies on the wire leaves the previously mirrored copy in place,
but back_finalize() returned without noting it, so the URL fell out of new.lst
and the end-of-update purge treated the file like a page that had vanished from
the site. Note the surviving copy, the way the incomplete-transfer branch above
already does, for the connection-level failures htsparse.c retries on. A
deliberate skip (too big, MIME-excluded, cancelled) keeps its current fate.

Closes #746

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-27 09:19:14 +02:00
Xavier Roche
6457828b72 tests: make the re-fetch assertions catch a crawl that never re-fetched
The new test passed with no second pass at all, so a regression that stopped
re-fetching would have looked green. Assert the failure the fixture provokes,
give stay.bin a fresh pass-2 body so a fix that stopped overwriting anything
fails, and compare reset.bin by checksum rather than by length.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-27 09:18:27 +02:00
Xavier Roche
f118cd86d3 tests: assert the failed re-fetch keeps its bytes on every platform
Test 93 filtered reset.bin out of its bucket lists because a connection killed
before the status line surfaced differently per platform. It no longer does, so
assert the resource like any other: unchanged, with the bytes pass 1 mirrored.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-27 09:05:56 +02:00
Xavier Roche
9fe47c3986 Remove dead HTS_USEZLIB guards now that zlib is mandatory (#761)
configure has rejected --without-zlib since #750, and htsglobal.h
#errors on a forced HTS_USEZLIB=0, so the macro can only ever be 1.
Collapse every #if/#ifdef HTS_USEZLIB guard (htsweb.c, htsname.c,
htslib.c, htszlib.c, htsselftest.c, htscodec.c) to its always-taken
branch; htsweb.c's guard used #ifdef where every other site used #if,
an inconsistency that no longer matters once the guard is gone.

htscodec.c's #else arms were a genuine zlib-free content-coding path
(Accept-Encoding: identity, hts_codec_unpack returning -1), not stubs.
Removed for consistency with the other nine guards: the cache
(htscache.c, proxy/store.c) already reaches minizip unconditionally
from ~60 call sites with no null backend, so a zlib-free build is not
actually reachable today regardless of this file.

No new test: this is dead-code removal with no behavior change.
Verified by differential build against master: htscodec.o and
htszlib.o are byte-identical; the other touched objects differ only
in __LINE__ immediates shifted by the removed guard lines (plus one
cosmetic objdump label-annotation artifact each in htsselftest.o and
htsweb.o, from string-literal pool reordering). Exported symbols in
libhttrack.so are unchanged. make check: 166/166 (158 pass, 8 skip).

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 09:05:52 +02:00
Xavier Roche
37fa549ac5 Add the regression test #747 could not carry when it landed (#760)
htsselftest.c and tests/Makefile.am were held by #718 while #747 was fixed, so
the thread-counting fix went in without a test. -#test=threadwait covers it
from both sides: a wait placed right after a spawn joins that thread, and
wait_n(n) leaves n running rather than draining them.

One spawn per round is what makes it bite. A batch gives the earlier threads
time to raise the counter themselves, which is why an eight-thread version
passed on the unfixed engine; one thread per round failed 10 runs out of 10.

The changes-race self-test can now drop the counter it kept because
htsthread_wait() could not be trusted to join.

Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 09:05:44 +02:00
Xavier Roche
1373530f17 A failed re-fetch overwrote the mirrored file with the aborted read's debris
A transfer that dies before a complete response has no body, yet the save path
still consulted r.adr, which at that point holds whatever the aborted header
read left behind: raw status-line bytes, or an empty buffer that truncated the
file to zero on macOS. Require a successful transfer, as the empty-body half of
the condition already did.

Closes #748

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-27 08:59:43 +02:00
Xavier Roche
de8c0eebfc Three hand-rolled copies of the unlink-then-rename fallback, with two return conventions (#754)
* One unlink-then-rename helper for the three copies

replace_file() (htsback.c), wacz_rename_over() (htswarc.c) and the inline
block in singlefile_rewrite_file() each worked around Windows' non-clobbering
rename, with two opposite return conventions and only one of the three
converting path separators. Copying the wrong one gets you an inverted success
test.

hts_rename_over() replaces all four call sites: hts_boolean return, fconv() on
both paths. It lands in htsname.c rather than the htstools.c the issue named,
to stay clear of an in-flight PR over that file.

The wacz test now runs a cacheless second pass over a poisoned copy of the
package the first pass wrote, so repackaging has to clobber an existing file.
43 and 74 also assert the failure warnings stay out of the log, which is all
an inverted return leaves behind at those two sites.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>

* rename helper: move hts_rename_over to htstools.c

Issue #726 asked for htstools.c; it went to htsname.c only because #718 held
the file at the time. Kept internal, not HTSEXT_API: htscore.h pulls in both
headers, so every call site reaches it either way.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* format: drop the trailing blank line left by the move

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <xroche@gmail.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 08:58:20 +02:00
Xavier Roche
9e29c1e159 Sitemap files are never read, so URLs nothing links to are never found (#718)
* Read sitemap files so URLs nothing links to are found

HTTrack finds URLs only by parsing links, so anything a site publishes solely
in its sitemap stayed invisible: robots.txt was already parsed, but its
Sitemap: lines were ignored and nothing else in the tree touched sitemaps.

Adds opt-in --sitemap (-%m), which probes the start host's robots.txt and
falls back to /sitemap.xml, and --sitemap-url (-%mu) for an explicit document.
Handles <urlset> and nested <sitemapindex>, plain or gzipped. Discovered URLs
enter with the full depth budget but still go through the wizard, so filters
and scope rules decide; a sitemap is not a filter bypass.

The parser reads attacker-controlled XML off the network, so it is capped on
URL count, index nesting, decompressed size and decompression ratio, and child
sitemaps must stay on the host that named them.

Closes #712

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* sitemap: distinguish the sitemapindex log line from a urlset one

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* sitemap: fix an out-of-bounds read, tighten the parser and the tests

lienrelatif() walked back from the last character of its current-path
argument without checking the path was non-empty, reading one byte before
the stack buffer. htsAddLink is the first caller to pass an empty savename,
which sitemap documents have because they are ingested rather than mirrored,
so ASan caught it on the new crawl test.

The parser drops a value whose numeric character reference decodes outside
printable ASCII, rather than leaving the reference verbatim and seeding a URL
the site never published, and classifies a document by its real root element,
so a comment naming the other one no longer flips urlset and sitemapindex.
The robots.txt line reader is bounded by the body size instead of relying on
a NUL terminator.

The self-test moves to 01_zlib-sitemap.test: MSan runs 01_engine-* only,
because an uninstrumented libz floods it with false positives.

Tests gain the assertions the earlier ones were missing: which of the
robots.txt route and the /sitemap.xml fallback was taken, that the sitemap
documents stay out of the mirror, that the off-host child sitemap is refused,
the sitemapindex nesting cap, the per-document URL cap at its production
value, and copy_htsopt coverage for the two new fields.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* sitemap: state what the decompression cap actually binds on

deflate tops out near 1032:1, so hts_codec_maxout never binds before the
64 MiB cap; the old comment implied a ratio guard that cannot fire.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* sitemap: a child sitemap is a fetch, so filters and robots.txt must gate it

An adversarial review found that a <sitemapindex> <loc>, and a robots.txt
Sitemap: line, went straight to hts_record_link: the request went out even
when a -* rule or robots.txt Disallow covered it. Only the <urlset> half ran
through the wizard. Gate the document itself on the filters and on
robots.txt, which is all that can apply: the wizard proper wants a referring
link, and its up/down travel rules would judge a child sitemap against the
parent sitemap's own directory. The robots.txt probe is exempt, being the
request that fetches the rules.

A 301 also used to end ingestion silently, since the engine re-queues the
target as a fresh link that carried no sitemap marking. That hit any site
redirecting http to https. The marking now follows the redirect.

The "N URL(s) added" counter reported what the scanner emitted rather than
what was taken, which hid both of the above; it now reads "N of M". The
fallback to /sitemap.xml keys on the same corrected count, so a robots.txt
whose only Sitemap: line is off-host or filtered still falls back. Root
classification skips a UTF-8 BOM and an XML namespace prefix, and the doc
list is cleared when a mirror starts rather than only when it ends.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* sitemap: gate each fetch by who asked for it, and anchor travel on the start URL

A nine-agent review found a scope escape: the sitemap document was its own
`premier`, so the wizard measured travel from wherever the site chose to put
its sitemap. A root /sitemap.xml therefore widened a /deep/dir/ crawl to the
whole host. The ingester now points the wizard at the crawl's own start link
and lets each seeded URL become its own anchor, which is what a command-line
seed gets.

Robots handling was both mistimed and undifferentiated. The Sitemap: lines are
now collected by robots_parse, on the same body in the same fetch, and acted on
after the parsed rules are installed rather than before; and the decision comes
from a new hts_robots_forbids extracted out of the wizard, so the sitemap path
inherits the -s1 filters-win override instead of a stricter hand-rolled check.
The four fetches are no longer treated alike: a sitemap the user names is user
intent, one the site declares invites the fetch, only the guessed /sitemap.xml
obeys a Disallow, and the URLs listed inside stay fully gated.

Also: hts_unescapeEntities replaces the private entity decoder, whose guard
tests and fuzz corpus it silently forfeited; hts_codec_head replaces hts_zhead,
which is only defined under HTS_USEZLIB; the composed URL buffer now fits two
maximal components plus a scheme, which a 2046-byte --sitemap-url reached; the
bounded search is promoted to htstools as hts_memstr; and the live state moves
from httrackp into htsoptstate, leaving two installed fields rather than three.

Tests gain the scope escape, the three robots cases, a cap-boundary control,
the handler invocation count and a compression-bomb decode. Every one was
checked against a deliberately broken build.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* sitemap: add the fuzz harness the parser was missing, and drop truncated Sitemap: lines

The file header called the scanner fuzzable while fuzz/ registered ten
harnesses and none for it. fuzz-sitemap feeds it raw XML, gzip-framed bodies
and truncated streams off a heap copy of exactly the input size, so an overread
is an ASan report rather than a quiet pass, with a four-file seed corpus.
60000 runs clean under ASan+UBSan.

robots_parse now drops a Sitemap: line that filled its scratch buffer instead
of handing on the half URL it was truncated to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* fuzz: keep only the four sitemap seed inputs

A libFuzzer run writes its finds into the first corpus directory, and 191 of
them were committed with the harness.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* selftest: bound the sitemap document builders' snprintf accumulation

snprintf returns the length it wanted to write, so accumulating it blind
lets the next offset and size argument walk past the buffer. Guard each
step the way the argv builder above already does, and give the per-URL
loop a real remaining-space bound instead of a fixed 33.

Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* sitemap: date the new files 2026

The headers were copied from an existing file and kept its 1998 year.

Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* sitemap: keep the ingestion state out of htsoptstate

htsoptstate is embedded by value as httrackp.state, so a field at its tail
shifts every httrackp member declared after it: an offsetof probe put
warc_file at 141752 on master and 141760 on the branch. Move the pointer to
httrackp's own tail, where every existing offset holds and copy_htsopt still
ignores it.

Also renumber the crawl test to 89, master having taken 87 and 90, and give
the new option8 checkbox the hidden companion that 90_webhttrack-checkbox-clear
requires, plus its row in that test's table.

Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: renumber the sitemap crawl test to 95

Master took 88 through 93 and #720 claims 94.

Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* sitemap: realign the htsopt.h comments after the single-file merge

Master's longer LLint declarator moved the block's comment column.

Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* sitemap: translate the new GUI strings into the remaining 28 locales (#738)

The feature PR added the four LANG_SITEMAP* entries to lang.def with English
and Francais only; every other locale fell back to English in the WebHTTrack
form. Each file is written in its own declared charset.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 07:59:59 +02:00
Xavier Roche
9571fb9a6a htsthread_wait() returned before threads it should join had started (#752)
process_chain was incremented by the child in hts_entry_point(), so a caller
that spawned threads and immediately called htsthread_wait() saw a zero count
and returned at once, free to tear down state the children were about to read.

Count at spawn instead, under the same mutex the waiter reads.

httrack.c freed the option block before waiting; wait first.

Closes #747

Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 07:57:47 +02:00
Xavier Roche
3e8595c46f zlib is mandatory: reject --without-zlib at configure instead of failing at link (#750)
--without-zlib only ever dropped -lz from LIBS; it never defined HTS_USEZLIB 0,
so every #if HTS_USEZLIB guard in the tree has been permanently true and the
build died with undefined references from minizip, htszlib.c, htswarc.c and
htsselftest.c. A zlib-free build is not reachable from there: the cache and the
WARC output are zip/gzip containers, and htsback.c already #error'd on
HTS_USEZLIB=0.

So make the requirement explicit. CHECK_ZLIB now errors out on --without-zlib
and on a missing header or library, keeping --with-zlib=DIR for a non-standard
prefix. htsback.c's #error moves to htsglobal.h where the knob is defined, with
a message that is accurate when it fires; its include of htszlib.h went with it,
unused.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 07:55:59 +02:00
29 changed files with 594 additions and 207 deletions

View File

@@ -46,10 +46,32 @@ jobs:
- name: Configure
run: |
set -euo pipefail
# Regenerate from configure.ac/Makefile.am to validate them; the
# committed generated files already let a plain checkout build.
# Regenerate: configure and the Makefile.in's are not tracked.
autoreconf -fi
# Disabling zlib must fail here rather than at link with a pile of
# undefined minizip references (#735). Both spellings, so a rewrite
# cannot keep one and lose the other. Probed out-of-tree to leave
# nothing behind.
nozlib="$RUNNER_TEMP/nozlib"
for arg in --without-zlib --with-zlib=no; do
rm -rf "$nozlib" && mkdir -p "$nozlib"
if (cd "$nozlib" && "$GITHUB_WORKSPACE/configure" "$arg" >out.log 2>&1); then
echo "::error::configure $arg succeeded; it must be rejected"
exit 1
fi
# ... and for the stated reason, not an unrelated configure failure.
grep -q "zlib cannot be disabled" "$nozlib/out.log" \
|| { cat "$nozlib/out.log"; exit 1; }
done
./configure
# Same dead end from the compile side. The bare compile is the
# control: without it a broken probe would pass vacuously.
hdr='#include "htsglobal.h"'
echo "$hdr" | $CC -I. -Isrc -fsyntax-only -xc -
if echo "$hdr" | $CC -DHTS_USEZLIB=0 -I. -Isrc -fsyntax-only -xc - 2>/dev/null; then
echo "::error::-DHTS_USEZLIB=0 compiled; htsglobal.h must reject it"
exit 1
fi
# a missing decoder would silently drop the coding from Accept-Encoding
grep -q "define HTS_USEBROTLI 1" config.h
grep -q "define HTS_USEZSTD 1" config.h

View File

@@ -175,7 +175,7 @@ AX_CHECK_ALIGNED_ACCESS_REQUIRED
# check for various headers
AC_CHECK_HEADERS([execinfo.h sys/ioctl.h])
### zlib
### zlib (mandatory)
CHECK_ZLIB()
### brotli and zstd content codings (optional)

View File

@@ -1,82 +1,33 @@
dnl @synopsis CHECK_ZLIB()
dnl
dnl This macro searches for an installed zlib library. If nothing
dnl was specified when calling configure, it searches first in /usr/local
dnl and then in /usr. If the --with-zlib=DIR is specified, it will try
dnl to find it in DIR/include/zlib.h and DIR/lib/libz.a. If --without-zlib
dnl is specified, the library is not searched at all.
dnl
dnl If either the header file (zlib.h) or the library (libz) is not
dnl found, the configuration exits on error, asking for a valid
dnl zlib installation directory or --without-zlib.
dnl
dnl The macro defines the symbol HAVE_LIBZ if the library is found. You should
dnl use autoheader to include a definition for this symbol in a config.h
dnl file. Sample usage in a C/C++ source is as follows:
dnl
dnl #ifdef HAVE_LIBZ
dnl #include <zlib.h>
dnl #endif /* HAVE_LIBZ */
dnl
dnl @version $Id$
dnl @author Loic Dachary <loic@senga.org>
dnl Look for zlib. It is a hard requirement, not an option: the cache and the
dnl WARC output are zip/gzip containers, and the bundled minizip calls zlib
dnl directly. --with-zlib=DIR points at a non-standard prefix.
dnl
dnl Adds -lz to LIBS and defines HAVE_LIBZ.
AC_DEFUN([CHECK_ZLIB],
#
# Handle user hints
#
[AC_MSG_CHECKING(if zlib is wanted)
AC_ARG_WITH(zlib,
[ --with-zlib=DIR root directory path of zlib installation [defaults to
/usr/local or /usr if not found in /usr/local]
--without-zlib to disable zlib usage completely],
[if test "$withval" != no ; then
AC_MSG_RESULT(yes)
ZLIB_HOME="$withval"
else
AC_MSG_RESULT(no)
fi], [
AC_MSG_RESULT(yes)
ZLIB_HOME=/usr/local
if test ! -f "${ZLIB_HOME}/include/zlib.h"
then
ZLIB_HOME=/usr
AC_DEFUN([CHECK_ZLIB], [
AC_ARG_WITH([zlib],
[AS_HELP_STRING([--with-zlib=DIR],[root directory of the zlib installation])],
[zlib_want=$withval], [zlib_want=yes])
if test "$zlib_want" = "no"; then
AC_MSG_ERROR([zlib cannot be disabled: the cache and the WARC output are zip/gzip containers, and the bundled minizip calls zlib directly])
fi
])
#
# Locate zlib, if wanted
#
if test -n "${ZLIB_HOME}"
then
ZLIB_OLD_LDFLAGS=$LDFLAGS
ZLIB_OLD_CPPFLAGS=$LDFLAGS
LDFLAGS="$LDFLAGS -L${ZLIB_HOME}/lib"
CPPFLAGS="$CPPFLAGS -I${ZLIB_HOME}/include"
AC_LANG_SAVE
AC_LANG_C
AC_CHECK_LIB(z, inflateEnd, [zlib_cv_libz=yes], [zlib_cv_libz=no])
AC_CHECK_HEADER(zlib.h, [zlib_cv_zlib_h=yes], [zlib_cv_zlib_h=no])
AC_LANG_RESTORE
if test "$zlib_cv_libz" = "yes" -a "$zlib_cv_zlib_h" = "yes"
then
#
# If both library and header were found, use them
#
AC_CHECK_LIB(z, inflateEnd)
AC_MSG_CHECKING(zlib in ${ZLIB_HOME})
AC_MSG_RESULT(ok)
else
#
# If either header or library was not found, revert and bomb
#
AC_MSG_CHECKING(zlib in ${ZLIB_HOME})
LDFLAGS="$ZLIB_OLD_LDFLAGS"
CPPFLAGS="$ZLIB_OLD_CPPFLAGS"
AC_MSG_RESULT(failed)
AC_MSG_ERROR(either specify a valid zlib installation with --with-zlib=DIR or disable zlib usage with --without-zlib)
fi
if test "$zlib_want" != "yes"; then
# An explicit prefix is authoritative: if the header is not under it,
# error rather than silently pick a system copy.
if test ! -f "$zlib_want/include/zlib.h"; then
AC_MSG_ERROR([zlib requested at $zlib_want but $zlib_want/include/zlib.h is missing])
fi
CPPFLAGS="$CPPFLAGS -I$zlib_want/include"
LDFLAGS="$LDFLAGS -L$zlib_want/lib"
elif test -f /usr/local/include/zlib.h; then
# Where the BSD ports tree lands zlib, and not always searched by default.
CPPFLAGS="$CPPFLAGS -I/usr/local/include"
LDFLAGS="$LDFLAGS -L/usr/local/lib"
fi
AC_CHECK_HEADER([zlib.h], [],
[AC_MSG_ERROR([zlib.h not found; install the zlib development files or pass --with-zlib=DIR])])
AC_CHECK_LIB([z], [inflateEnd], [],
[AC_MSG_ERROR([libz not found; install the zlib development files or pass --with-zlib=DIR])])
])

View File

@@ -48,11 +48,6 @@ Please visit our Website: http://www.httrack.com
#include "htsftp.h"
#include "htscodec.h"
#include "htsproxy.h"
#if HTS_USEZLIB
#include "htszlib.h"
#else
#error HTS_USEZLIB not defined
#endif
#ifdef _WIN32
#ifndef __cplusplus
@@ -590,12 +585,17 @@ static int create_back_tmpfile(httrackp *opt, lien_back *const back,
return 0;
}
/* Move src onto dst; RENAME does not clobber an existing target on Windows. */
static hts_boolean replace_file(const char *src, const char *dst) {
if (RENAME(src, dst) == 0)
return HTS_TRUE;
(void) UNLINK(dst);
return RENAME(src, dst) == 0 ? HTS_TRUE : HTS_FALSE;
/* Did the fetch fail to produce a response, as opposed to the engine
deliberately passing the resource over? Only the latter may be purged. */
static hts_boolean back_transfer_failed(const int statuscode) {
switch (statuscode) {
case STATUSCODE_TOO_BIG:
case STATUSCODE_EXCLUDED:
case STATUSCODE_TEST_OK:
return HTS_FALSE;
default:
return statuscode <= 0 ? HTS_TRUE : HTS_FALSE;
}
}
/* Commit or restore a re-fetch backup (#77 follow-up): a re-fetch over an
@@ -615,7 +615,7 @@ static void back_finalize_backup(httrackp *opt, lien_back *const back,
}
/* On failure keep the backup: an orphaned temp beats losing the good copy.
*/
if (!replace_file(back->tmpfile, back->url_sav))
if (!hts_rename_over(back->tmpfile, back->url_sav))
hts_log_print(opt, LOG_WARNING | LOG_ERRNO,
"could not restore %s; previous copy kept as %s",
back->url_sav, back->tmpfile);
@@ -746,7 +746,7 @@ int back_finalize(httrackp * opt, cache_back * cache, struct_back * sback,
"Read error when decompressing");
}
UNLINK(unpacked);
} else if (replace_file(unpacked, back[p].url_sav)) {
} else if (hts_rename_over(unpacked, back[p].url_sav)) {
/* The temp bypassed filecreate(), which is what chmods. */
#ifndef _WIN32
chmod(back[p].url_sav, HTS_ACCESS_FILE);
@@ -1044,6 +1044,14 @@ int back_finalize(httrackp * opt, cache_back * cache, struct_back * sback,
/* Aborted, error, or not ready: url_sav (if written) is broken; restore the
previous copy from the backup. */
back_finalize_backup(opt, &back[p], HTS_FALSE);
/* Note the surviving copy, or the end-of-update purge drops what this run
never managed to replace (#746). */
if (!back[p].testmode && back_transfer_failed(back[p].r.statuscode) &&
back[p].url_sav[0] != '\0' && fexist_utf8(back[p].url_sav)) {
filenote(&opt->state.strc, back[p].url_sav, NULL);
file_notify(opt, back[p].url_adr, back[p].url_fil, back[p].url_sav, 0, 0,
back[p].r.notmodified);
}
return -1;
}

View File

@@ -72,7 +72,7 @@ hts_codec hts_codec_parse(const char *encoding) {
return HTS_CODEC_IDENTITY;
if (strfield2(encoding, "gzip") || strfield2(encoding, "x-gzip") ||
strfield2(encoding, "deflate") || strfield2(encoding, "x-deflate"))
return HTS_USEZLIB ? HTS_CODEC_DEFLATE : HTS_CODEC_UNSUPPORTED;
return HTS_CODEC_DEFLATE;
if (strfield2(encoding, "br"))
return HTS_USEBROTLI ? HTS_CODEC_BROTLI : HTS_CODEC_UNSUPPORTED;
if (strfield2(encoding, "zstd"))
@@ -98,16 +98,11 @@ hts_codec hts_codec_parse(const char *encoding) {
const char *hts_acceptencoding(hts_boolean compressible, hts_boolean secure) {
if (!compressible)
return "identity";
#if HTS_USEZLIB
/* br and zstd over TLS only, as browsers do: a cleartext intermediary that
rewrites a coding it can not read would corrupt the mirror. */
if (secure)
return "gzip, deflate" HTS_AE_BROTLI HTS_AE_ZSTD ", identity;q=0.9";
return "gzip, deflate, identity;q=0.9";
#else
(void) secure;
return "identity";
#endif
}
hts_boolean hts_codec_is_archive_ext(hts_codec codec, const char *ext) {
@@ -300,11 +295,7 @@ int hts_codec_unpack(hts_codec codec, const char *filename,
return -1;
switch (codec) {
case HTS_CODEC_DEFLATE:
#if HTS_USEZLIB
return hts_zunpack(filename, newfile);
#else
return -1;
#endif
case HTS_CODEC_BROTLI:
case HTS_CODEC_ZSTD:
break;
@@ -348,10 +339,8 @@ size_t hts_codec_head(hts_codec codec, const void *in, size_t in_len, void *out,
if (in == NULL || in_len == 0 || out == NULL || out_len == 0)
return 0;
switch (codec) {
#if HTS_USEZLIB
case HTS_CODEC_DEFLATE:
return hts_zhead(in, in_len, out, out_len);
#endif
#if HTS_USEBROTLI
case HTS_CODEC_BROTLI:
return codec_head_brotli(in, in_len, out, out_len);

View File

@@ -1980,10 +1980,9 @@ int httpmirror(char *url1, httrackp * opt) {
}
// ATTENTION C'EST ICI QU'ON SAUVE LE FICHIER!!
// An empty body must not overwrite the file when the transfer failed
// (statuscode <= 0, e.g. an -M hard-stop): it would truncate a good
// copy to 0 (#77 follow-up).
if (r.adr != NULL || (r.size == 0 && r.statuscode > 0)) {
// A failed transfer has no body: r.adr holds debris from the aborted
// read, which would destroy the copy being re-fetched (#748).
if (r.statuscode > 0 && (r.adr != NULL || r.size == 0)) {
file_notify(opt, urladr(), urlfil(), savename(), 1, 1, r.notmodified);
if (filesave(opt, r.adr, (int) r.size, savename(), urladr(), urlfil()) !=
0) {
@@ -2693,7 +2692,11 @@ HTSEXT_API int structcheck(const char *path) {
if (!S_ISDIR(st.st_mode)) {
#if HTS_REMOVE_ANNOYING_INDEX
if (S_ISREG(st.st_mode)) { /* Regular file in place ; move it and create directory */
sprintf(tmpbuf, "%s.txt", file);
/* bounded here, not by the path-length guard far above */
if (!sprintfbuff(tmpbuf, "%s.txt", file)) {
errno = ENAMETOOLONG;
return -1;
}
if (rename(file, tmpbuf) != 0) { /* Can't rename regular file */
return -1;
}
@@ -2801,7 +2804,11 @@ HTSEXT_API int structcheck_utf8(const char *path) {
if (!S_ISDIR(st.st_mode)) {
#if HTS_REMOVE_ANNOYING_INDEX
if (S_ISREG(st.st_mode)) { /* Regular file in place ; move it and create directory */
sprintf(tmpbuf, "%s.txt", file);
/* bounded here, not by the path-length guard far above */
if (!sprintfbuff(tmpbuf, "%s.txt", file)) {
errno = ENAMETOOLONG;
return -1;
}
if (RENAME(file, tmpbuf) != 0) { /* Can't rename regular file */
return -1;
}
@@ -3821,11 +3828,12 @@ int htsAddLink(htsmoduleStruct * str, char *link) {
strcpybuff(codebase, heap(ptr)->fil);
else
strcpybuff(codebase, heap(heap(ptr)->precedent)->fil);
a = codebase + strlen(codebase) - 1;
// empty codebase has no last char; codebase-1 would underflow
a = codebase[0] != '\0' ? codebase + strlen(codebase) - 1 : codebase;
while((*a) && (*a != '/') && (a > codebase))
a--;
if (*a == '/')
*(a + 1) = '\0'; // couper
*(a + 1) = '\0'; // cut
} else { // couper http:// éventuel
if (strfield(codebase, "http://")) {
char BIGSTK tempo[HTS_URLMAXSIZE * 2];

View File

@@ -138,10 +138,11 @@ Please visit our Website: http://www.httrack.com
#define HTS_DOSNAME 0
#endif
// utiliser zlib?
// zlib is mandatory: the cache is a zip and minizip calls it regardless
#ifndef HTS_USEZLIB
// autoload
#define HTS_USEZLIB 1
#elif !HTS_USEZLIB
#error HTS_USEZLIB=0 is not a supported configuration
#endif
// brotli and zstd content codings; off unless the build opted in (configure,

View File

@@ -1131,12 +1131,10 @@ int http_sendhead(httrackp * opt, t_cookie * cookie, int mode,
// Compression accepted ?
if (retour->req.http11) {
hts_boolean compressible = HTS_FALSE;
hts_boolean compressible =
(!retour->req.range_used && !retour->req.nocompression);
hts_boolean secure = HTS_FALSE;
#if HTS_USEZLIB
compressible = (!retour->req.range_used && !retour->req.nocompression);
#endif
#if HTS_USEOPENSSL
secure = retour->ssl ? HTS_TRUE : HTS_FALSE;
#endif

View File

@@ -43,9 +43,7 @@ Please visit our Website: http://www.httrack.com
#include "htsencoding.h"
#include "htssniff.h"
#include "htscodec.h"
#if HTS_USEZLIB
#include "htszlib.h"
#endif
#include <ctype.h>
#include <limits.h>

View File

@@ -43,6 +43,7 @@ Please visit our Website: http://www.httrack.com
#include "htsglobal.h"
#include "htscore.h"
#include "htsmodules.h"
#include "htsback.h"
#include "htsdefines.h"
#include "htslib.h"
@@ -63,9 +64,7 @@ Please visit our Website: http://www.httrack.com
#include "htswarc.h"
#include "htschanges.h"
#include "htssinglefile.h"
#if HTS_USEZLIB
#include "htszlib.h"
#endif
#if HTS_USEZSTD
#include <zstd.h>
#endif
@@ -2623,6 +2622,68 @@ static int st_savename(httrackp *opt, int argc, char **argv) {
return 0;
}
/* an empty fil started htsAddLink's codebase walk before the buffer (#730) */
static int st_addlink(httrackp *opt, int argc, char **argv) {
htsmoduleStruct BIGSTK str;
cache_back cache;
struct_back *sback;
hash_struct hash;
int ptr = 0;
int i;
(void) argc;
(void) argv;
memset(&cache, 0, sizeof(cache));
cache.hashtable = (void *) coucal_new(0);
sback = back_new(opt, opt->maxsoc * 32 + 1024);
/* same wiring as hts_mirror (htscore.c) */
hash_init(opt, &hash, opt->urlhack);
hash.liens = (const lien_url *const *const *) &opt->liens;
opt->hash = &hash;
hts_record_init(opt);
memset(&str, 0, sizeof(str));
str.opt = opt;
str.sback = sback;
str.cache = &cache;
str.hashptr = &hash;
str.ptr_ = &ptr;
str.addLink = htsAddLink;
/* [0] is the underflow; [1] and [2] are controls that the trim is unchanged.
A query-only link is the one that notices the trim at all: for the others
ident_url_relatif() re-derives the directory from the path it is given. */
for (i = 0; i < 3; i++) {
static const char *const fil[3] = {"", "/dir/page.html", "/dir/page.html"};
static const char *const lnk[3] = {"sub/page.html", "sub/page.html",
"?x=1"};
static const char *const want[3] = {
"untouched", "http://www.example.com/dir/sub/page.html",
"http://www.example.com/dir/?x=1"};
char BIGSTK loc[HTS_URLMAXSIZE * 2];
char BIGSTK link[HTS_URLMAXSIZE];
strcpybuff(loc, "untouched");
strcpybuff(link, lnk[i]);
str.localLink = loc;
str.localLinkSize = (int) sizeof(loc);
if (!hts_record_link(opt, "www.example.com", fil[i], "", "", "", ""))
return 1;
ptr = heap_top_index();
str.url_host = heap(ptr)->adr;
str.url_file = heap(ptr)->fil;
assertf(htsAddLink(&str, link) == 0); /* refused by the wizard either way */
if (strcmp(loc, want[i]) != 0) {
fprintf(stderr, "addlink[%d]: got '%s' want '%s'\n", i, loc, want[i]);
return 1;
}
}
printf("addlink self-test OK\n");
return 0;
}
static int st_cache(httrackp *opt, int argc, char **argv) {
int err;
@@ -3264,6 +3325,81 @@ static int st_topindex(httrackp *opt, int argc, char **argv) {
return 0;
}
/* Build a path of exactly len chars under base; returns that length. */
static size_t st_structcheck_longpath(char *dst, size_t dstsize,
const char *base, size_t len) {
size_t n = strlen(base);
assertf(len < dstsize && n + 2 <= len);
memmove(dst, base, n);
while (n < len) {
size_t seg = len - n - 1;
if (seg > 200) /* stay under the usual 255-byte component limit */
seg = len - n == 202 ? 199 : 200; /* never leave a bare separator */
dst[n++] = '/';
memset(dst + n, 'x', seg);
n += seg;
}
dst[n] = '\0';
return n;
}
/* The path guard, and the <name>.txt rename structcheck() performs when a
regular file sits where a directory has to go (#745). */
static int st_structcheck(httrackp *opt, int argc, char **argv) {
char BIGSTK path[HTS_URLMAXSIZE * 2];
char BIGSTK target[HTS_URLMAXSIZE * 2];
FILE *fp;
(void) opt;
if (argc < 1) {
fprintf(stderr, "usage: -#test=structcheck <writable directory>\n");
return 1;
}
/* over the guard: refused before a single directory is created */
st_structcheck_longpath(path, sizeof(path), argv[0], HTS_URLMAXSIZE + 1);
errno = 0;
assertf(structcheck(path) == -1);
assertf(errno == EINVAL);
errno = 0;
assertf(structcheck_utf8(path) == -1);
assertf(errno == EINVAL);
{
char *const sep = strchr(path + strlen(argv[0]) + 1, '/');
assertf(sep != NULL);
sep[1] = '\0'; /* the outermost component it would have created */
assertf(!dir_exists(path));
}
/* a regular file where a directory belongs is renamed to <name>.txt */
snprintf(path, sizeof(path), "%s/sc", argv[0]);
fp = fopen(path, "wb");
assertf(fp != NULL);
fclose(fp);
snprintf(path, sizeof(path), "%s/sc/sub/", argv[0]);
assertf(structcheck(path) == 0);
assertf(dir_exists(path));
snprintf(target, sizeof(target), "%s/sc.txt", argv[0]);
assertf(fexist(target));
/* the utf-8 entry point carries the same rename */
snprintf(path, sizeof(path), "%s/u8", argv[0]);
fp = FOPEN(path, "wb");
assertf(fp != NULL);
fclose(fp);
snprintf(path, sizeof(path), "%s/u8/sub/", argv[0]);
assertf(structcheck_utf8(path) == 0);
assertf(dir_exists(path));
snprintf(target, sizeof(target), "%s/u8.txt", argv[0]);
assertf(fexist_utf8(target));
printf("structcheck self-test OK\n");
return 0;
}
/* Each inplace_escape_*() must equal escape_*() on a copy. */
static int st_inplace_escape(httrackp *opt, int argc, char **argv) {
/* >255 bytes forces the helper's malloct path, not the stack buffer */
@@ -3377,7 +3513,6 @@ static int st_status(httrackp *opt, int argc, char **argv) {
return 0;
}
#if HTS_USEZLIB
/* Deflate src->path at windowBits (16+ gzip, + zlib, - raw); 0 on success. */
static int ae_write_packed(const char *path, int windowBits,
const unsigned char *src, size_t len) {
@@ -3451,7 +3586,6 @@ static int ae_write_collision(const char *path, const unsigned char *src,
freet(buf);
return ok ? 0 : 1;
}
#endif
/* Write src[0..len) to path as-is; 0 on success. */
static int ae_write_raw(const char *path, const unsigned char *src,
@@ -3501,7 +3635,6 @@ static int st_acceptencoding(httrackp *opt, int argc, char **argv) {
assertf(strstr(on, "br") == NULL && strstr(on, "zstd") == NULL);
assertf((strstr(tls, ", br") != NULL) == (HTS_USEBROTLI != 0));
assertf((strstr(tls, "zstd") != NULL) == (HTS_USEZSTD != 0));
#if HTS_USEZLIB
if (argc >= 1) {
static const int windowBits[] = {16 + MAX_WBITS, MAX_WBITS, -MAX_WBITS};
const unsigned char small[] =
@@ -3580,10 +3713,6 @@ static int st_acceptencoding(httrackp *opt, int argc, char **argv) {
}
freet(body);
}
#else
(void) argc;
(void) argv;
#endif
printf("acceptencoding self-test OK: %s\n", on);
return 0;
}
@@ -3928,7 +4057,6 @@ static int st_sitemap(httrackp *opt, int argc, char **argv) {
freet(big);
}
#if HTS_USEZLIB
/* A highly compressible document decodes without running away: the ratio
budget cannot bind (deflate tops out near 1032:1), so this pins the
decompression path itself rather than the 64 MiB ceiling. */
@@ -3979,13 +4107,11 @@ static int st_sitemap(httrackp *opt, int argc, char **argv) {
freet(z);
freet(x);
}
#endif
/* An unterminated <loc> at end of buffer must not read past it. */
assertf(sm_scan("<urlset><loc>http://h.test/a", 100, &idx, &c) == 0);
assertf(sm_scan("<urlset><lo", 100, &idx, &c) == 0);
#if HTS_USEZLIB
/* A gzip-framed document is decompressed before scanning. */
{
const char *const xml =
@@ -4015,7 +4141,6 @@ static int st_sitemap(httrackp *opt, int argc, char **argv) {
assertf(hts_sitemap_scan(z, 4, 100, &idx, sm_take, &c) == -1);
freet(z);
}
#endif
/* robots.txt: only Sitemap: records, comments stripped, case-insensitive,
and group-independent (no User-agent line needed). */
@@ -6089,6 +6214,116 @@ static int st_changes(httrackp *opt, int argc, char **argv) {
return err;
}
/* #747: a thread is outstanding from the moment hts_newthread() returns, not
from the moment it starts running, or a wait right after the spawn joins
nothing. One thread per round is what makes the old bug visible: the wait
had to find the counter at zero, and a batch of spawns gives the earlier
threads time to raise it. Unfixed, one round in two caught it, so the round
count is what turns that into a reliable failure. */
#define THREADWAIT_N 8
#define THREADWAIT_ROUNDS 16
#define THREADWAIT_SLEEP_MS 50
#define THREADWAIT_GATE_MS 10000
static htsmutex threadwait_lock = HTSMUTEX_INIT;
static int threadwait_done = 0;
static hts_boolean threadwait_gated = HTS_FALSE;
static int threadwait_count(void) {
int n;
hts_mutexlock(&threadwait_lock);
n = threadwait_done;
hts_mutexrelease(&threadwait_lock);
return n;
}
static void threadwait_thread(void *arg) {
(void) arg;
Sleep(THREADWAIT_SLEEP_MS);
hts_mutexlock(&threadwait_lock);
threadwait_done++;
hts_mutexrelease(&threadwait_lock);
}
/* Stays outstanding until the gate clears, so wait_n() can be asked to leave a
known number of live threads behind. Bounded: a wait_n() that wrongly drains
them would otherwise never return, and hang the suite instead of failing. */
static void threadwait_gated_thread(void *arg) {
int waited;
(void) arg;
for (waited = 0; waited < THREADWAIT_GATE_MS; waited += 10) {
hts_boolean gated;
hts_mutexlock(&threadwait_lock);
gated = threadwait_gated;
hts_mutexrelease(&threadwait_lock);
if (!gated)
break;
Sleep(10);
}
hts_mutexlock(&threadwait_lock);
threadwait_done++;
hts_mutexrelease(&threadwait_lock);
}
static int st_threadwait(httrackp *opt, int argc, char **argv) {
int err = 0;
int i, round;
(void) opt;
(void) argc;
(void) argv;
/* htsthread_wait() joins a thread spawned just before it */
for (round = 0; round < THREADWAIT_ROUNDS && !err; round++) {
hts_mutexlock(&threadwait_lock);
threadwait_done = 0;
hts_mutexrelease(&threadwait_lock);
if (hts_newthread(threadwait_thread, NULL) != 0) {
fprintf(stderr, "threadwait: cannot spawn\n");
return 1;
}
htsthread_wait();
if (threadwait_count() != 1) {
fprintf(stderr, "threadwait: round %d returned before the thread ran\n",
round);
err = 1;
}
}
/* htsthread_wait_n(n) leaves n behind rather than draining everything */
hts_mutexlock(&threadwait_lock);
threadwait_done = 0;
threadwait_gated = HTS_TRUE;
hts_mutexrelease(&threadwait_lock);
for (i = 0; i < THREADWAIT_N; i++) {
if (hts_newthread(threadwait_gated_thread, NULL) != 0) {
fprintf(stderr, "threadwait: cannot spawn a gated thread\n");
return 1;
}
}
htsthread_wait_n(THREADWAIT_N);
if (threadwait_count() != 0) {
fprintf(stderr, "threadwait: wait_n(%d) joined %d gated threads\n",
THREADWAIT_N, threadwait_count());
err = 1;
}
hts_mutexlock(&threadwait_lock);
threadwait_gated = HTS_FALSE;
hts_mutexrelease(&threadwait_lock);
htsthread_wait();
if (threadwait_count() != THREADWAIT_N) {
fprintf(stderr, "threadwait: wait left %d/%d gated threads running\n",
THREADWAIT_N - threadwait_count(), THREADWAIT_N);
err = 1;
}
printf("threadwait self-test: %s\n", err ? "FAIL" : "OK");
return err;
}
#define CHANGES_RACE_FILES 8
#define CHANGES_RACE_ROUNDS 400
@@ -6103,11 +6338,8 @@ static void changes_race_notify(httrackp *opt, int n) {
hts_changes_notify(opt, "race.example", fil, save, HTS_TRUE, HTS_FALSE);
}
/* htsthread_wait() counts a thread only once it is running, so it can return
before any of them started; join on our own counter instead. */
static htsmutex changes_race_lock = HTSMUTEX_INIT;
static int changes_race_started = 0;
static int changes_race_live = 0;
static int changes_race_count(int *which) {
int n;
@@ -6127,9 +6359,6 @@ static void changes_race_thread(void *arg) {
hts_mutexrelease(&changes_race_lock);
for (i = 0; i < CHANGES_RACE_ROUNDS; i++)
changes_race_notify(opt, i % CHANGES_RACE_FILES);
hts_mutexlock(&changes_race_lock);
changes_race_live--;
hts_mutexrelease(&changes_race_lock);
}
/* A transfer thread the crawl never joins (FTP) reaches hts_changes_notify()
@@ -6184,7 +6413,6 @@ static int st_changes_race(httrackp *opt, int argc, char **argv) {
threads reaching a fresh one together race on the init itself. */
hts_mutexlock(&changes_race_lock);
changes_race_started = 0;
changes_race_live = 4;
hts_mutexrelease(&changes_race_lock);
for (i = 0; i < 4; i++) {
if (hts_newthread(changes_race_thread, opt) != 0) {
@@ -6198,8 +6426,7 @@ static int st_changes_race(httrackp *opt, int argc, char **argv) {
for (i = 0; i < 64; i++)
hts_changes_report(opt, &out);
hts_changes_close_opt(opt);
while (changes_race_count(&changes_race_live) > 0)
Sleep(10);
htsthread_wait();
/* Sealed: a straggler must be dropped, not start a report nobody writes. */
changes_race_notify(opt, CHANGES_RACE_FILES + 1);
@@ -6277,6 +6504,8 @@ static const struct selftest_entry {
st_changes},
{"changes-race", "<dir>", "--changes under a late transfer thread (#714)",
st_changes_race},
{"threadwait", "", "htsthread_wait() joins threads spawned just before it",
st_threadwait},
{"pause", "", "randomized inter-file pause target self-test", st_pause},
{"relative", "<link> <curr-file>", "relative link between two paths",
st_relative},
@@ -6307,6 +6536,8 @@ static const struct selftest_entry {
{"fsize", "<dir>", "file size past the 2GB signed-32-bit wrap", st_fsize},
{"growsize", "", "buffer capacity for a 64-bit file size (no int wrap)",
st_growsize},
{"addlink", "", "htsAddLink codebase walk over an empty current path",
st_addlink},
{"cache", "<dir>", "cache read/write round-trip self-test", st_cache},
{"cacheindex", "", "cache-index (.ndx) parse must stay in bounds",
st_cacheindex},
@@ -6330,6 +6561,9 @@ static const struct selftest_entry {
{"useragent", "", "default User-Agent self-test", st_useragent},
{"makeindex", "[dir]", "hts_finish_makeindex footer/refresh self-test",
st_makeindex},
{"structcheck", "<dir>",
"structcheck path guard and the <name>.txt rename it performs",
st_structcheck},
{"topindex", "[dir]",
"hts_buildtopindex charset handling of a non-ASCII project dir",
st_topindex},

View File

@@ -1108,13 +1108,8 @@ hts_boolean singlefile_rewrite_file(httrackp *opt, const char *root,
(void) chmod(fconv(catbuff, sizeof(catbuff), StringBuff(tmp)),
HTS_ACCESS_FILE);
#endif
if (ok) {
/* RENAME does not clobber an existing target on Windows. */
if (RENAME(StringBuff(tmp), page_path) != 0) {
(void) UNLINK(page_path);
ok = RENAME(StringBuff(tmp), page_path) == 0 ? HTS_TRUE : HTS_FALSE;
}
}
if (ok)
ok = hts_rename_over(StringBuff(tmp), page_path);
if (!ok) {
hts_log_print(opt, LOG_ERROR, "single-file: could not rewrite %s",
page_path);

View File

@@ -44,9 +44,18 @@ Please visit our Website: http://www.httrack.com
#endif
#endif
/* Outstanding threads, counted at spawn rather than by the child at entry, so
that a caller which spawns and immediately waits still joins them (#747). */
static int process_chain = 0;
static htsmutex process_chain_mutex = HTSMUTEX_INIT;
static void process_chain_add(int delta) {
hts_mutexlock(&process_chain_mutex);
process_chain += delta;
assertf(process_chain >= 0);
hts_mutexrelease(&process_chain_mutex);
}
HTSEXT_API void htsthread_wait(void) {
htsthread_wait_n(0);
}
@@ -99,20 +108,12 @@ static void *hts_entry_point(void *tharg)
void *const arg = s_args->arg;
void (*fun) (void *arg) = s_args->fun;
free(tharg);
hts_mutexlock(&process_chain_mutex);
process_chain++;
assertf(process_chain > 0);
hts_mutexrelease(&process_chain_mutex);
freet(tharg);
/* run */
fun(arg);
hts_mutexlock(&process_chain_mutex);
process_chain--;
assertf(process_chain >= 0);
hts_mutexrelease(&process_chain_mutex);
process_chain_add(-1);
#ifdef _WIN32
return 0;
#else
@@ -122,18 +123,20 @@ static void *hts_entry_point(void *tharg)
/* create a thread */
HTSEXT_API int hts_newthread(void (*fun) (void *arg), void *arg) {
hts_thread_s *s_args = malloc(sizeof(hts_thread_s));
hts_thread_s *s_args = malloct(sizeof(hts_thread_s));
assertf(s_args != NULL);
s_args->arg = arg;
s_args->fun = fun;
process_chain_add(1);
#ifdef _WIN32
{
unsigned int idt;
HANDLE handle =
(HANDLE) _beginthreadex(NULL, 0, hts_entry_point, s_args, 0, &idt);
if (handle == 0) {
free(s_args);
process_chain_add(-1);
freet(s_args);
return -1;
} else {
/* detach the thread from the main process so that is can be independent */
@@ -151,7 +154,8 @@ HTSEXT_API int hts_newthread(void (*fun) (void *arg), void *arg) {
|| pthread_attr_setstacksize(&attr, stackSize) != 0
|| (retcode =
pthread_create(&handle, &attr, hts_entry_point, s_args)) != 0) {
free(s_args);
process_chain_add(-1);
freet(s_args);
return -1;
} else {
/* detach the thread from the main process so that is can be independent */

View File

@@ -1443,3 +1443,15 @@ HTSEXT_API hts_boolean hts_findissystem(find_handle find) {
}
return 0;
}
hts_boolean hts_rename_over(const char *src, const char *dst) {
char csrc[CATBUFF_SIZE], cdst[CATBUFF_SIZE];
fconv(csrc, sizeof(csrc), src);
fconv(cdst, sizeof(cdst), dst);
if (RENAME(csrc, cdst) == 0)
return HTS_TRUE;
/* RENAME does not clobber an existing target on Windows. */
(void) UNLINK(cdst);
return RENAME(csrc, cdst) == 0 ? HTS_TRUE : HTS_FALSE;
}

View File

@@ -137,6 +137,10 @@ HTSEXT_API hts_boolean hts_findisdir(find_handle find);
HTSEXT_API hts_boolean hts_findisfile(find_handle find);
HTSEXT_API hts_boolean hts_findissystem(find_handle find);
/* Move src onto dst, replacing an existing dst; HTS_TRUE on success. Both
paths are fconv()'d. */
hts_boolean hts_rename_over(const char *src, const char *dst);
#endif
#endif

View File

@@ -964,16 +964,6 @@ static zipFile wacz_zip_open(const char *path) {
return zipOpen2_64(path, 0 /*create*/, NULL, &ff);
}
/* Move src onto dst (UTF-8/Windows-safe); RENAME won't clobber on Windows, so
fall back to unlink+rename. Returns 0 on success. */
static int wacz_rename_over(const char *src, const char *dst) {
char cs[CATBUFF_SIZE], cd[CATBUFF_SIZE];
if (RENAME(fconv(cs, sizeof(cs), src), fconv(cd, sizeof(cd), dst)) == 0)
return 0;
(void) UNLINK(fconv(cd, sizeof(cd), dst));
return RENAME(fconv(cs, sizeof(cs), src), fconv(cd, sizeof(cd), dst));
}
/* Package the segment(s) + .cdx + a generated pages.jsonl into <base>.wacz at
crawl end (the archive file(s) and .cdx are already closed on disk). */
static void warc_wacz_package(warc_writer *w) {
@@ -1092,7 +1082,7 @@ static void warc_wacz_package(warc_writer *w) {
hts_log_print(w->opt, LOG_WARNING,
"WACZ: packaging failed, kept existing %s untouched",
waczpath);
} else if (wacz_rename_over(tmppath, waczpath) != 0) {
} else if (!hts_rename_over(tmppath, waczpath)) {
(void) UNLINK(fconv(catbuff, sizeof(catbuff), tmppath));
hts_log_print(w->opt, LOG_WARNING | LOG_ERRNO,
"WACZ: could not finalize %s", waczpath);

View File

@@ -101,8 +101,8 @@ static void htsweb_sig_brpipe(int code) {
/* ignore */
}
/* Number of background threads */
static int background_threads = 0;
/* Threads that never return; no wait may count on them draining. */
static int nonjoinable_threads = 0;
/* Server/client ping handling */
static htsmutex pingMutex = HTSMUTEX_INIT;
@@ -224,9 +224,7 @@ int main(int argc, char *argv[]) {
#ifdef HTS_USESWF
smallserver_setkey("USESWF", "1");
#endif
#ifdef HTS_USEZLIB
smallserver_setkey("USEZLIB", "1");
#endif
#ifdef _WIN32
smallserver_setkey("WIN32", "1");
#endif
@@ -299,15 +297,19 @@ int main(int argc, char *argv[]) {
/* pinger */
if (parentPid > 0) {
hts_newthread(client_ping, (void *) (uintptr_t) parentPid);
background_threads++; /* Do not wait for this thread! */
if (hts_newthread(client_ping, (void *) (uintptr_t) parentPid) == 0) {
#ifndef _WIN32
nonjoinable_threads++; /* client_ping() only ever leaves through exit() */
#endif
}
smallserver_setpinghandler(pingHandler, NULL);
}
/* launch */
ret = help_server(argv[1], defaultPort, bindAddr);
htsthread_wait_n(background_threads - 1);
/* Drain everything a mirror may still have in flight, the pinger aside. */
htsthread_wait_n(nonjoinable_threads);
hts_uninit();
#ifdef _WIN32
@@ -382,7 +384,6 @@ void webhttrack_main(char *cmd) {
commandRunning = 1;
DEBUG(fprintf(stderr, "commandRunning=1\n"));
hts_newthread(back_launch_cmd, (void *) strdup(cmd));
background_threads++; /* Do not wait for this thread! */
}
void webhttrack_lock(void) {
@@ -423,8 +424,8 @@ static int webhttrack_runmain(httrackp * opt, int argc, char **argv) {
/* Rock'in! */
ret = hts_main2(argc, argv, opt);
/* Wait for pending threads to finish */
htsthread_wait_n(background_threads);
/* Wait for pending threads to finish; the pinger and this thread stay. */
htsthread_wait_n(nonjoinable_threads + 1);
return ret;
}

View File

@@ -40,7 +40,6 @@ Please visit our Website: http://www.httrack.com
#include "htscodec.h"
#include "htszlib.h"
#if HTS_USEZLIB
/* zlib */
/*
#include <zlib.h>
@@ -274,4 +273,3 @@ const char *hts_get_zerror(int err) {
break;
}
}
#endif

View File

@@ -306,8 +306,8 @@ int main(int argc, char **argv) {
fprintf(stderr, "* %s\n", hts_errmsg(opt));
}
global_opt = NULL;
htsthread_wait(); /* pending threads still read opt */
hts_free_opt(opt);
htsthread_wait(); /* wait for pending threads */
hts_uninit();
#ifdef _WIN32

View File

@@ -0,0 +1,7 @@
#!/bin/bash
#
# an empty current path underflowed htsAddLink's codebase walk (#730).
set -euo pipefail
httrack -O /dev/null -#test=addlink | grep -q "addlink self-test OK"

View File

@@ -0,0 +1,12 @@
#!/bin/bash
#
# structcheck(): the path-length guard, and the <name>.txt rename that once
# built its target with an unbounded sprintf (#745).
set -euo pipefail
dir=$(mktemp -d)
trap 'rm -rf "$dir"' EXIT
httrack -O /dev/null -#test=structcheck "$dir" |
grep -q "structcheck self-test OK"

View File

@@ -0,0 +1,28 @@
#!/bin/bash
#
set -euo pipefail
tmpdir=$(mktemp -d "${TMPDIR:-/tmp}/httrack_threadwait_st.XXXXXX") || exit 1
trap 'rm -rf "$tmpdir"' EXIT HUP INT QUIT PIPE TERM
# No pipe into grep: SIGPIPE would mask a failing exit status.
expect_ok() {
local label="$1" out
shift
out=$("$@" 2>&1) || {
echo "FAIL: ${label} exited non-zero: ${out}"
exit 1
}
case "$out" in
*"${label}: OK"*) ;;
*)
echo "FAIL: ${out}"
exit 1
;;
esac
}
# A thread is outstanding from the moment hts_newthread() returns, so a wait
# that follows the spawn joins it, and wait_n(n) still leaves n behind (#747).
expect_ok "threadwait self-test" httrack -O "${tmpdir}/o1" -#test=threadwait

View File

@@ -18,6 +18,7 @@ set -euo pipefail
bash "$top_srcdir/tests/local-crawl.sh" \
--rerun-args '--update -M400000' \
--log-found 'More than 400000 bytes have been transferred.. giving up' \
--log-not-found 'could not restore' \
--found bigtrunc/slow.bin \
--file-min-bytes bigtrunc/slow.bin 655360 \
--file-min-bytes bigtrunc/fast.bin 655360 \

View File

@@ -79,6 +79,29 @@ done
# 65535 is a valid port the old "< 65535" bound refused
web_accepted 65535
# --- htsserver returns from main() instead of blocking at exit -------------
# The exit wait must exclude exactly the threads that never return (#753).
# $tmp has no lang.def, so the server fails right after announcing and main()
# reaches the wait on its own; rc, not EXITED, carries the verdict.
web_exits() {
local rc=0
: >"$tmp/exit.log"
run_with_timeout 30 htsserver "$tmp" "$@" >"$tmp/exit.log" 2>&1 || rc=$?
grep -q "^EXITED" "$tmp/exit.log" ||
! echo "FAIL: #753: htsserver ${*:-(no options)} never finished serving" || exit 1
# exactly 1, not merely "not 124": a crash or an assertf abort also escapes
# the wait, and would otherwise read as a pass
test "$rc" -eq 1 ||
! echo "FAIL: #753: htsserver ${*:-(no options)} exited $rc, want 1" || exit 1
}
# no pinger: the excluded count went negative
web_exits
# a pinger, which never returns and so must stay excluded
web_exits --ppid $$
# --- proxytrack <proxy-addr:port> <ICP-addr:port> --------------------------
# A bad argument falls through to the usage screen; it had no range check at
# all, so 65616 quietly listened on port 80. A valid one binds and blocks.

View File

@@ -5,11 +5,16 @@
# real gate: STORE-mode entries, the fixed layout, recomputed sha256 digests and
# the datapackage-digest chain (plus py-wacz/pywb when importable). The crawl
# skips cleanly on a build without OpenSSL (no conformant SHA-256 -> no package).
#
# The cacheless second pass repackages over the .wacz the first one left: the
# clobber, not the happy rename, is what breaks when a platform refuses to
# rename onto an existing file (#726).
set -eu
: "${top_srcdir:=..}"
bash "$top_srcdir/tests/local-crawl.sh" --errors 0 --wacz-validate \
--rerun-args '-C0' --log-not-found 'could not finalize' \
--found 'mini304/index.html' --found 'mini304/page.html' \
httrack 'BASEURL/mini304/index.html' --warc-file warc-out --wacz

View File

@@ -94,10 +94,7 @@ expect_counts_match() {
echo "OK"
}
lines() { printf '%s\n' "$@" | sort; }
# A connection killed before the status line surfaces differently per platform
# (macOS truncates the file, #748; Linux leaves it), so reset.bin is asserted on
# its own and kept out of the exact lists.
listed_but_reset() { listed "$1" | grep -v '/reset\.bin$' | sort; }
digest() { cksum <"$1" | tr -d '[:space:]'; }
# --- pass 1: nothing to compare against --------------------------------------
httrack "${common[@]}" -O "$out" --purge-old=0 "${base}/changes/index.html" \
@@ -120,6 +117,7 @@ grep -aq "first crawl" "${out}/hts-log.txt" || {
}
host="127.0.0.1_${port}"
resetdigest=$(digest "${out}/${host}/changes/reset.bin")
# A leftover from some earlier failed attempt, at a name pass 2 fetches for the
# first time: on disk, but never part of the previous mirror, so it is new.
printf 'leftover junk' >"${out}/${host}/changes/fresh.html"
@@ -144,21 +142,19 @@ expect "gone is doomed.html and the old redirect target" \
expect "changed is exactly the four moved resources" \
"$(lines "${host}/changes/coded.bin" "${host}/changes/index.html" \
"${host}/changes/moved.bin" "${host}/changes/moved.html")" \
"$(listed_but_reset changed)"
"$(listed changed | sort)"
# Re-served byte for byte, 200, no Last-Modified, no ETag. flaky.bin gets there
# through a failed transfer and a retry, so its file is notified twice;
# redirtarget.html arrives behind a 302. sized.html has a byte-identical
# payload, but the rewritten link to the renamed redirect target makes the file
# on disk change length.
expect "unchanged is exactly the six stable resources" \
# on disk change length. reset.bin never completes a transfer (#746), so its
# previous copy stands.
expect "unchanged is exactly the seven stable resources" \
"$(lines "${host}/changes/codedstable.bin" "${host}/changes/flaky.bin" \
"${host}/changes/redirtarget.html" "${host}/changes/sized.html" \
"${host}/changes/stable.bin" "${host}/changes/stable.html")" \
"$(listed_but_reset unchanged)"
# reset.bin never completes a transfer, so it drops out of new.lst with its
# mirrored copy still there. Whatever else it is, it is not a deletion.
expect "a failed re-fetch is never reported gone" "" \
"$(listed gone | grep '/reset\.bin$' || true)"
"${host}/changes/redirtarget.html" "${host}/changes/reset.bin" \
"${host}/changes/sized.html" "${host}/changes/stable.bin" \
"${host}/changes/stable.html")" \
"$(listed unchanged | sort)"
expect_counts_match "pass 2 counts match its lists"
# Every mirrored file appears once and only once, across all four buckets.
@@ -192,19 +188,15 @@ test -f "${out}/${host}/changes/doomed.html" || {
exit 1
}
echo "OK"
printf '[a failed transfer keeps its file] ..\t'
test -f "${out}/${host}/changes/reset.bin" || {
echo "FAIL: reset.bin lost its previous copy"
exit 1
}
echo "OK"
expect "a failed transfer keeps its bytes" "$resetdigest" \
"$(digest "${out}/${host}/changes/reset.bin")"
# --- pass 3: the same run with purging on ------------------------------------
httrack "${common[@]}" -O "$out" --update "${base}/changes/index.html" \
>"${tmpdir}/log3" 2>&1
expect "the report says files were purged" "true" "$(field purged)"
expect "gone is the page that only pass 2 linked" \
"${host}/changes/transient.html" "$(listed_but_reset gone)"
"${host}/changes/transient.html" "$(listed gone | sort)"
printf '[a purged file is off disk] ..\t'
test ! -f "${out}/${host}/changes/transient.html" || {
echo "FAIL: transient.html survived --purge-old"

View File

@@ -0,0 +1,25 @@
#!/bin/bash
#
# A re-fetch cut mid-header must leave the mirrored file alone (#748) and keep
# it out of the update purge (#746). err.bin is the control, an HTTP 500 on the
# same resource already surviving both; stay.bin answers normally with a new
# body, so a fix that stopped overwriting anything fails.
set -euo pipefail
: "${top_srcdir:=..}"
bash "$top_srcdir/tests/local-crawl.sh" \
--rerun-args '--update' \
--log-found 'after 2 retries at link .*keep/page\.html' \
--log-found 'after 2 retries at link .*keep/data\.bin' \
--file-matches 'keep/page.html' 'KEEP-PAGE-V1' \
--file-not-matches 'keep/page.html' 'HTTP/1\.0 200' \
--file-matches 'keep/data.bin' 'KEEP-BIN-V1' \
--file-not-matches 'keep/data.bin' 'HTTP/1\.0 200' \
--file-min-bytes 'keep/data.bin' 2048 \
--file-matches 'keep/err.bin' 'KEEP-ERR-V1' \
--file-min-bytes 'keep/err.bin' 2048 \
--file-matches 'keep/stay.bin' 'KEEP-STAY-V2' \
--file-min-bytes 'keep/stay.bin' 2048 \
httrack 'BASEURL/keep/index.html' --retries=2

View File

@@ -30,6 +30,7 @@ TEST_EXTENSIONS = .test
TEST_LOG_COMPILER = $(BASH)
TESTS = \
00_runnable.test \
01_engine-addlink.test \
01_engine-changes.test \
01_engine-charset.test \
01_engine-cmdline.test \
@@ -81,7 +82,9 @@ TESTS = \
01_engine-status.test \
01_engine-stripquery.test \
01_engine-strsafe.test \
01_engine-structcheck.test \
01_engine-syscharset.test \
01_engine-threadwait.test \
01_engine-topindex.test \
01_engine-urlhack.test \
01_engine-unescape-bounds.test \
@@ -194,6 +197,7 @@ TESTS = \
92_local-proxytrack-ndx-fields.test \
93_local-changes.test \
94_local-single-file.test \
95_local-sitemap.test
95_local-sitemap.test \
96_local-refetch-keep.test
CLEANFILES = check-network_sh.cache

View File

@@ -57,6 +57,8 @@ rerun_dead=
tmpdir=
serverpid=
crawlpid=
wacz_poisoned=
wacz_poison="stale-wacz-that-a-second-pass-must-replace"
function warning {
echo "** $*" >&2
@@ -272,6 +274,14 @@ if test -n "$warc_validate"; then
test -z "$w1" || cp "$w1" "${tmpdir}/warc-pass1.gz"
fi
# Poison the first-pass .wacz when a second pass follows: repackaging moves the
# new archive over it, so the marker must be gone afterwards (#726). Poisoning
# beats comparing the two packages, which can come out byte-identical.
if test -n "$wacz_validate" && test -n "${rerun}${rerun_args}"; then
wacz_poisoned=$(find "$mirrorroot" -maxdepth 2 -name '*.wacz' 2>/dev/null | sort | tail -n1)
test -z "$wacz_poisoned" || echo "$wacz_poison" >"$wacz_poisoned"
fi
# --- optional second pass: re-mirror into the same dir (cache/update path) ----
if test -n "$rerun"; then
info "re-running httrack (update pass)"
@@ -416,6 +426,12 @@ if test -n "$wacz_validate"; then
fi
die "no .wacz file produced under $mirrorroot"
fi
if test -n "$wacz_poisoned"; then
info "checking the second pass replaced the .wacz"
grep -q "$wacz_poison" "$wacz_poisoned" 2>/dev/null &&
die "stale .wacz kept: $wacz_poisoned"
result "OK"
fi
validator=$(nativepath "${testdir}/wacz-validate.py")
info "validating WACZ package"
"$python" "$validator" "$(nativepath "$wacz")" >&2 || die "WACZ validation failed"

View File

@@ -1000,6 +1000,62 @@ class Handler(SimpleHTTPRequestHandler):
v = 1 if self.refetch_pass() == 1 else 2
self.send_raw(b"<html><body><p>STAY-V%d</p></body></html>" % v, "text/html")
# --- re-fetch cut mid-header, so nothing is stored (#746, #748) ---------
KEEP_PAGE = b"<html><body><p>KEEP-PAGE-V1</p></body></html>"
KEEP_BIN = b"KEEP-BIN-V1\n" + b"\x51\x52\x53\x54" * 512
KEEP_ERR = b"KEEP-ERR-V1\n" + b"\x61\x62\x63\x64" * 512
def send_cut_headers(self):
"""Hang up mid-header: the response never becomes parseable."""
self.close_connection = True
try:
self.wfile.write(b"HTTP/1.0 200 OK\r\nContent-Ty")
self.wfile.flush()
except OSError:
pass
self.connection.close()
def route_keep_index(self):
self.refetch_pass()
self.send_html(
'\t<a href="page.html">page</a>\n'
'\t<a href="data.bin">data</a>\n'
'\t<a href="err.bin">err</a>\n'
'\t<a href="stay.bin">stay</a>\n'
)
def route_keep_page(self):
if self.refetch_pass() == 1:
self.send_raw(self.KEEP_PAGE, "text/html")
else:
self.send_cut_headers()
def route_keep_data(self):
if self.refetch_pass() == 1:
self.send_raw(self.KEEP_BIN, "application/octet-stream")
else:
self.send_cut_headers()
# Control: an HTTP error on the same resource already keeps the copy, so
# the cut-header routes above must end up indistinguishable from it.
def route_keep_err(self):
if self.refetch_pass() == 1:
self.send_raw(self.KEEP_ERR, "application/octet-stream")
else:
self.send_response(500)
self.send_header("Content-Type", "text/html")
self.send_header("Content-Length", "0")
self.end_headers()
# Control: answers normally, with a new body on pass 2, so a fix that
# stopped overwriting mirrored files altogether would be caught.
def route_keep_stay(self):
v = 1 if self.refetch_pass() == 1 else 2
self.send_raw(
b"KEEP-STAY-V%d\n" % v + b"\x71\x72\x73\x74" * 512,
"application/octet-stream",
)
# Echo what httrack advertised, so a crawl can assert the header.
def route_codec_ae(self):
self.send_raw(
@@ -1950,6 +2006,11 @@ class Handler(SimpleHTTPRequestHandler):
"/uptrunc/page.html": route_uptrunc_page,
"/uptrunc/file.bin": route_uptrunc_file,
"/uptrunc/stay.html": route_uptrunc_stay,
"/keep/index.html": route_keep_index,
"/keep/page.html": route_keep_page,
"/keep/data.bin": route_keep_data,
"/keep/err.bin": route_keep_err,
"/keep/stay.bin": route_keep_stay,
"/types/index.html": route_types_index,
"/types/control.php": route_types,
"/types/photo.png": route_types,