* Two installed headers declare symbols the library does not export
htsbasenet.h declares openssl_ctx and htsnet.h declares
hts_dns_set_resolver_backend, both hidden by -fvisibility=hidden and both
installed into $(includedir)/httrack. A consumer that includes either header
and uses the name compiles, then fails to link.
Move both behind HTS_INTERNAL_BYTECODE, the guard the other internal
declarations in the installed set already use. Every in-tree caller defines
it, so the engine build is unaffected.
Test 207 installs the headers, derives the hidden set from the library's
symbol table minus its dynamic table, and links a probe for every hidden name
a consumer can reach through an installed header.
Closes#977
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Test 207's floor against a vacuous run counts only header-static symbols
Every one of the 31 symbols that reaches the link probe is a static function
defined in the installed header itself (abortf_ and the rest of htssafe.h,
StringOom_ from htsstrings.h). The probe object compiles its own copy, so the
link cannot fail whatever the library exports, and "probed >= 5" holds without
the harvest ever seeing an extern declaration. Cutting the header loop down to
htssafe.h alone leaves probed at 7, and the test still passes with 13 of the 14
headers unchecked.
So plant a leak: a canary header declaring the hidden symbol the negative
control has already shown cannot link. The loop has to report it, which puts
the install, the harvest, the intersection and the link on the same path a real
leak takes. The cut-down loop now fails.
The HTTRACK_SHLIB comment also blamed the skip on static-only builds. macOS
skips as well, where libtool names the library .dylib.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Trim test 207's comments to the house one-line default
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Declare struct addrinfo at file scope in htsnet.h
Without it the tag in the resolver-backend prototypes is a fresh type scoped
to its own declaration wherever <netdb.h> has not already declared it.
Closes#987
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Crash report can hang: the symbolizer forks from the signal handler
fork() runs the pthread_atfork prepare handlers before it forks, and glibc's
malloc registers one that takes every arena lock. A signal raised inside
malloc (glibc's own heap-corruption detector aborts from exactly there, and
SIGABRT is wired to sig_fatal) then leaves the handler blocked on a lock its
own thread holds, and the report is lost.
vfork() bypasses the atfork handlers, at the price of a child that may only
issue syscalls: the PATH search execvp() performs allocates, so the symbolizer
is resolved once at startup and the child execv()s an absolute path.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Test 182 pins a 64-bit address width, so it fails the i386 leg
addr2line -a zero-pads the address to the target's pointer size: 8 nibbles
against a 32-bit ELF, 16 against a 64-bit one. The gate is "is addr2line
installed", not the architecture, so the {16} match fails a correct build on
the gcc -m32 leg and on Debian's 32-bit buildds. Test 80 already matches the
same line unanchored.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Spawn the symbolizer with posix_spawnp() instead of fork()
posix_spawnp() reaches the same child through clone(CLONE_VM|CLONE_VFORK),
which runs no pthread_atfork handler, so the arena lock a signal raised inside
malloc already holds cannot deadlock the report away.
It also does the PATH search itself and reports an exec failure through its
return value, so the hand-rolled find_on_path(), the two statics caching its
answer and the BT_NO_SYMBOLIZER exit protocol all go. Building the file actions
in hts_backtrace_init() keeps the only allocation off the crash path, and with
the rewrite happening in the parent the addr2line to llvm-symbolizer fallback
comes back.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Assert the symbolizer output lands on the report fd
The child's stdout redirect is now a file actions object built at startup, so a
wrong target sends symbolized frames to the crawler's own stdout instead. The
existing checks merge both streams and cannot see that.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Take the report fd out of hts_print_backtrace()'s signature
The child's stdout redirect is prebuilt at init against stderr, so passing any
other fd quietly skipped symbolization on a runtime check. The one caller
passed stderr, and htsbacktrace.h ships with the program and not with the
installed dev headers, so there is no ABI cost to making the wrong value
impossible to write.
Test 182 also picks up test 80's ARM gate. A build with no unwind tables
prints "unwinding failed" instead of the OS notice, which 182's skip check
misses, so it would have gone red on armhf rather than skipping.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Assert a symbol name in test 182, not just the frame shape
An all-"??" symbolization regression still emits the address lines.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* escape_remove_control() leaves the original tail glued to the result
It compacted the non-control bytes downward in place but never wrote the
new terminator, so any input it actually shortened came back as the
compacted head plus whatever the old tail left behind: "/a\013bc" came
out "/abcc". The function is HTSEXT_API and declared in the installed
httrack-library.h, so the broken contract is a public one.
The parser can reach it: VT and FF pass the link scan as is_space()
members and the later strip only removes CR, LF and TAB, so a page with
href="/a<VT>bc" fetched /abcc before this.
Closes#974
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* escape-control self-test lets a shifted cut and a stray write through
Grading mutants against the new test found two survivors. No vector held
a space, so moving the loop's cut to `c > 32` stripped it and the suite
stayed green. Nothing reached above 0x0c either, so `c >= 13` passed too.
The added vectors pin both sides of the cut, and the second drops two
bytes instead of one.
The canary read one byte at `inlen + 1`, so a stray write two or more
past the compacted end went unseen. The buffer is already poisoned in
full, so scanning the rest of it costs nothing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Release 3.49.16
Version bump across configure.ac, htsglobal.h, version.rc and the AppStream
metainfo, plus the release notes for history.txt and debian/changelog.
VERSION_INFO 3:7:0 -> 3:8:0: no installed struct or exported signature moved
this cycle, so the soname stays .so.3 and Debian needs no package rename.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Correct the 3.49.16 release notes
The AppStream metainfo describes WebHTTrack and is read by Linux software
centres, so it now lists only WebHTTrack-visible changes: the macOS bullet and
the nine option-dialog strings (WinHTTrack-only, no hits outside lang/) are out,
the charset re-encoding is in. Drop the task-switcher claim, which webhttrack
cannot make: it is a script that launches a browser.
Fix two over-claims in history.txt. Only the trailer section after the
terminating chunk was rejected, not any chunked page. The cache aborted the
mirror on an over-long URL rather than merely refusing one it had accepted.
Cite #901 rather than the PR that closed it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The macOS leg builds tests/altstackprobe.c like every other platform, and it
has neither off64_t nor mmap64. SYS_gettid is Linux-only too, so gating just
the glibc half would still not compile there.
The whole tracer now lives behind __linux__, with mmap64 behind __GLIBC__
inside it: musl has no LFS split and its plain mmap() is already the one the
engine calls. Nothing outside the gate changed, so on macOS the file is the
one that was already building, plus four headers it has. The trace is what
183's second leg reads, and that leg is Linux-only by the same reasoning.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
183 could only see the install: crash_stack_thread never returns, so
hts_entry_point never reaches the leave hook and every release-side bug was
structurally invisible. altstackprobe grows an LD_PRELOAD tracer for
sigaltstack(), mmap() and munmap(), and a second leg drives -#test=threadwait,
whose workers do return: 25 stacks installed, 24 handed back, each unmap right
behind its own SS_DISABLE, and the main thread's kept.
Also: hts_set_thread_hooks refuses half a pair, crash_threadstack aborts on a
spawn failure instead of leaving the test to blame the handler, and the -#c
kinds list clips rather than aborts now the table has outgrown its 64 bytes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Installed headers do not compile under -std=c99
htssafe.h spells typeof without the underscores in three GNU-only
branches and calls POSIX strnlen from its inline bodies. __GNUC__
survives -std=c99 but neither the keyword nor the declaration does, so
htssafe.h and the three installed headers including it stop a strict-ISO
consumer.
Use __typeof__, and route the inline helpers through a private wrapper
that calls the libc strnlen wherever it is actually declared. tests/206
compiles every installed header under -std=c99 and -std=c11.
Closes#972
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Test 206 asks for more POSIX than it needs
The two socket headers get -D_POSIX_C_SOURCE so getnameinfo() and
NI_NUMERICHOST resolve. At 200809L that also un-hides strnlen, so
htsnet.h and htsopt.h stopped covering the strnlen half of #972: delete
the __STRICT_ANSI__ guard from htssafe_strnlen_ and both still compile
clean. 200112L exposes the networking surface and leaves strnlen hidden,
so both catch that mutant again.
The level is probed rather than hard-coded, because only glibc is
verified here. A libc that needs 200809L still passes, and the test
reports the lost coverage.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Test 206 never expands the macros it is meant to protect
Two of the three __typeof__ sites this PR fixes sit in
htsbuff_must_be_array_, which is a macro. A macro body only reaches the
compiler where it expands, and 206 compiles a bare #include of each
header, so reverting either of those two lines to plain typeof left the
test green. Only HTS_IS_CHAR_BUFFER was covered, and then only because
htsarrays.h happens to expand it.
Add a translation unit that instantiates htsbuff_array(), htsbuff_catn()
and strcpybuff(). Reverting any of the three sites now reds 206; before
this, only one of them did.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Test 206 never runs the byte loop it protects
206 is -fsyntax-only throughout, so htssafe_strnlen_'s fallback was covered
for parsing and nothing else: changing its bound to `i <= maxlen` or its
result to `i - 1` left the test green. That loop is the strict-ISO
consumer's strnlen, and this header bounds copies from the wire.
Compile and run one strict-ISO TU that differentials htssafe_strnlen_
against memchr over every content/bound pair up to 8 bytes, on a poisoned
buffer so an overshot bound cannot land on a NUL and read as correct. The
gate is asserted first, or a build that fell back to libc would grade libc.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Fail rather than silently drop strnlen coverage in test 206
The 200809L fallback un-hides strnlen, so it would report PASS while covering
less than it claims. No CI libc reaches it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Mark hts_record_assert_memory_failed HTS_UNUSED (#1002)
A C++ consumer of the installed headers hits -Wunused-function on this
static function: it is only called from a TypedArray macro expansion,
so a translation unit that just includes htsarrays.h leaves it unused.
g++ only runs that check on a real compile, not -fsyntax-only, which
is why 206's C loop (all -fsyntax-only) missed it; the new C++ leg
uses a real -c compile for exactly this reason. htssafe.h already
tags the same shape HTS_UNUSED.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
lang/Slovak.txt line 166 read Hľdanie where the windows-1250-encoded
Slovak word is Hľadanie, missing the "a". Pre-existing typo left over
from #963's charset-declaration move, unrelated to that PR's changes.
Closes#994
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The frame match keys on addr2line's " at " separator, which binutils
translates: under fr_FR.UTF-8 it reads " a" and the test fails on a
report whose frames are all present.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Nothing checks that the data directory comes from the OS path, not argv[0]
Reverting the startup lines to argv[0] alone left the whole suite green: the
part -#test=datadir cannot see is which path gets handed to
hts_resolve_datadir(). 215_engine-datadir-ospath runs a copy of the engine from
a fabricated bin/ under two lying argv[0]s, one naming a decoy tree that has
templates of its own and one a bare name as a PATH lookup leaves, and asserts
the templates the run resolved are the copy's.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Fall back to src/httrack when the build needed no libtool wrapper
--disable-shared leaves the program in src/ with no .libs copy, and the test
then failed on a missing HTTRACK_BIN instead of running.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Plant a binary at the decoy path so an existence-checked argv[0] is caught
The decoy tree had templates but no program, so an engine that prefers argv[0] only when it names an existing file reached the OS path by accident and passed all three runs. Gating that preference on access(argv[0], F_OK) reproduces it: the test as merged passes that engine, this one fails it on the decoy run.
The decoy also goes first on PATH for the bare-name run, so an engine resolving argv[0] by lookup has something wrong to find.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Make six lang/*.txt files agree with their own LANGUAGE_CHARSET
Croatian, Slovak, Polski and Slovenian are stored in CP1250 and Macedonian in
CP1251, but they declared an ISO-8859 charset. html/server/*.html copies
LANGUAGE_CHARSET into its <meta> and htsserver.c decodes the POST body with it,
so every s-caron, z-caron and Cyrillic letter came out wrong in webhttrack.
Strings added since #588 followed the declaration instead of the surrounding
bytes, leaving the files mixed.
The declaration moves to the codepage the bytes are actually in, and the minority
lines are transcoded to match. Going the other way was not an option: Latin-1
cannot spell Slovenian at all, and ISO-8859-2 would lose the typographic quotes
and ellipsis Polski uses. windows-1250 and windows-1251 are already in the
engine's codepage table and already declared by Cesky, Russian, Bulgarian,
Ukrainian and Uzbek.
Svenska declared ISO-8859-2 over Latin-1 bytes, which agree on everything it uses
except a-ring; only the declaration changes there.
62_lang-integrity now decodes every file through its declared charset and rejects
both a byte the charset has no mapping for and a decode landing in the C1 range,
which is the tell of a Windows codepage wearing an ISO-8859 label. Two
deliberately mislabelled copies guard the check itself, one per branch.
Closes#963
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Keep the charset guard honest on macOS
BSD sed does not read \r in an s/// replacement, so the C1 probe would have
been labelled ISO-8859-2r there and failed on the unknown charset name instead
of on the C1 branch it exists to cover. The extractor strips CR anyway, so the
escape bought nothing. Also default n to 1, since a grep that errors leaves it
empty and the comparison then reads as a pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Name the executable's own frames in a crash report
dladdr() reports the main program's path as argv[0], which is a bare name
when the binary was found on PATH. The symbolizer's access() check then
fails and it drops the whole module, so every frame of the executable stays
a raw offset. A --disable-shared build puts the entire engine there, which
is why 105_suite-timeout failed and 80_engine-crash-symbolize skipped.
Resolve /proc/self/exe once at init and match frames against the main
program's load base, so the symbolizer gets a path it can open.
Not a link-flag problem: relinking with -rdynamic exports none of these
symbols, since -fvisibility=hidden and file-scope statics make them
STB_LOCAL and --export-dynamic only promotes global ones.
Closes#889
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Reject a clipped /proc/self/exe path
readlink() cannot report truncation, so an executable path of BT_PATH_SIZE
bytes or more came back as a prefix that find_main_object() kept. The prefix
fits copy_bounded(), and access() accepts it whenever it names a readable
directory, so the crash report gained a 1023-byte header pointing at that
directory and addr2line's "is a directory" complaint in place of the
executable's frames. A full buffer now reads as unresolved, which falls back
to argv[0] as before.
Checked by hand: an httrack installed 1024 bytes deep and run off PATH prints
the raw trace only, and 1023 still symbolizes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Pin test 157 to one executable block and the symbol column
A second block for the executable means library frames were misattributed to
it, and the bare symbol match could hit a path component named main.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
sigaltstack() state is per-thread, so the alternate stack #866 gave the main
thread does nothing for the engine's workers, and the crawl recurses there.
A worker running out of stack leaves the kernel no room for a signal frame,
and it is killed outright with no report.
htsthread gains a pair of hooks the thread start routine runs around the
worker body, and httrack registers the altstack installer and its release.
Releasing matters now that it is per-thread: 64kB of mapping per worker,
never given back, adds up over a crawl.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Add a weekly networked AppStream metainfo check
The per-PR lint job and tests/210 validate with --no-net so an httrack.com
outage never reds an unrelated PR, but that also hides the one failure a
software centre user would notice: a dead screenshot URL. Add a scheduled
workflow that runs appstreamcli validate without --no-net weekly, and on
demand via workflow_dispatch; it does not touch the existing --no-net gates.
Closes#971
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Tighten the appstream-network comments to one line each
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
fuzz/Makefile.am's noinst_PROGRAMS only builds under --enable-fuzzers, so #960
left fuzz/ alone. An anchored per-binary list keeps the tracked fuzz-*.c
sources out of the pattern, unlike a fuzz-* glob that would swallow them too.
Closes#970
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* A relocated install cannot find its own shared library
httrack and htsserver now link with a loader-relative rpath,
$ORIGIN/../lib (@loader_path/../lib on macOS), beside the absolute
libdir libtool records, so a tree configured for one prefix and
unpacked somewhere else finds libhttrack instead of failing before
main.
Not every loader expands that token, so configure probes it instead of
trusting the linker: it builds a small shared library and an executable
that needs it, then runs the executable from an unrelated directory
with the library path cleared. Solaris and the older BSDs get a second
attempt with -Wl,-z,origin. Nothing is added when libdir is already on
the loader's search path, which covers every distribution package.
Closes#906
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Put the configured libdir ahead of the binary-relative rpath
Emitting $ORIGIN/../lib first shadowed the configured libdir whenever
$(bindir)/../lib is a different directory, which a --libdir=$prefix/lib64
or a split-bindir layout makes routine: a stale library sitting there
won the lookup on an ordinary, non-relocated install. Naming $(libdir)
ourselves before the token puts it back in front, so the token only
answers once the configured libdir is gone.
The system-libdir gate no longer reads libtool's
sys_lib_dlsearch_path_spec, which is only augmented when /etc/ld.so.conf
exists and would let a distribution libdir through in a sysroot or a
minimal container; it pattern-matches /lib, /lib64, /usr/lib and
/usr/lib64 and their subdirectories instead. /usr/local is no longer
treated as a system libdir, so a default configure now enables the
rpath and the tests exercise it.
Cross compiling no longer enables the feature unprobed. It defaults to
off with --enable-origin-rpath as the override, and that override takes
the -Wl,-z,origin spelling Solaris and the older BSDs need, since
nothing can run to tell the two apart.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Cover the rpath ordering and the cases that must not get one
195_install-relocate.test asserted only that the token was present, so
it could not see it landing ahead of the configured libdir, and it
skipped outright wherever the rpath was suppressed. It now reads the
rpath entries in order and requires the libdir to come first, requires
proxytrack to carry no token at all, and in a suppressed build asserts
the absence rather than skipping.
Proving the relocated copy started was also not enough: a libhttrack
elsewhere on the loader path would have answered. The copy's own library
is now moved away and the binary has to stop working, which replaces the
ldconfig guard (found by `command -v ldconfig`, which fails for a user
without /sbin on PATH, and matching any soname including a .so.2 that
could never satisfy .so.3) and works on macOS, where there is no ldd.
Staging through DESTDIR rather than overriding the paths keeps libtool
from rewriting the build tree while the rest of the suite runs.
196_install-rpath-gates.test configures a system prefix, a multiarch
libdir, --disable-origin-rpath and --disable-shared and asserts each
one emits no rpath, with a private prefix first so a probe that never
succeeds cannot make the rest vacuous. One shared cache keeps the five
runs to a few seconds.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Suppress the rpath for the 32-bit multilib libdirs too
/lib32, /libx32 and their /usr counterparts are on the loader path, and
libtool's sys_lib_dlsearch_path_spec used to cover them, so replacing
that read with an enumerated list handed a 32-bit distribution build an
rpath it never had and lintian's binary-or-shlib-defines-rpath with it.
Still enumerated rather than globbed as /lib*, which would swallow
/libfoo, /usr/library and /usr/libexec.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Configure a symlink farm, not the already-configured srcdir
196_install-rpath-gates.test configures out of $abs_top_srcdir, which
automake rejects when a config.status sits there. CI builds in-tree, so
every leg that runs the full suite failed; the out-of-tree local runs and
the msan leg (restricted TESTS) never saw it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Refresh mergeability after master moved
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Turn the binary-relative rpath off on Darwin, and read Mach-O with otool
dyld consults LC_RPATH only for an @rpath/ load path, and libtool stamps
libhttrack with an absolute -install_name, so the @loader_path entry the
Darwin arm emitted was never read. The copied tree stayed just as broken.
Gate the feature off there instead of shipping a dead load command.
That left the two bugs the macOS leg was actually failing on. Test 195
picked its object dumper by whatever happened to be installed, and an ELF
reader handed a Mach-O prints no rpath rather than an error, so a present
rpath read as an absent one; it now picks by platform. Test 196's nested
configures inherit none of the parent's CPPFLAGS/LDFLAGS, so brew's
keg-only openssl turned the default --enable-https into a hard configure
error. The gate has nothing to do with TLS, so they ask for =auto.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* A stack-overflow SIGSEGV kills httrack with no diagnostic at all
sig_fatal went in with plain signal(), so it ran on the faulting thread's
own stack. When the fault is stack exhaustion there is no room left for a
signal frame, the kernel falls back to the default action, and the process
dies at 139 having printed nothing. Register the four fatal signals through
sigaction() with SA_ONSTACK, over a 64 kB alternate stack allocated at
startup, and prime backtrace() at init so glibc's lazy libgcc_s dlopen()
does not happen inside the handler.
Only the main thread is covered: htsbacktrace.c is linked into the httrack
binary rather than into libhttrack, so the engine's worker threads cannot
reach the installer. They behave as before, SA_ONSTACK being ignored where
no alternate stack exists.
Closes#866
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Test 180 could not tell a stack overflow from an ordinary fault
Three mutants passed it: a crash_stack() that only dereferences NULL, a
backtrace() truncated to one frame, and a backtrace that returns nothing at
all, since the frame assertion accepted "No stack trace available" as a
match. The reports now have to differ: a runaway recursion repeats the same
return address ~250 times where the segv control repeats none, and the
control asserts the threshold still separates them.
Also stop skipping on macOS and the BSDs, where sigaltstack() exists and the
fix is live; only the trace text is gated on Linux. Each assertion now names
what failed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* The alt stack leaked, and took over the sanitizer's
LeakSanitizer called the 64 kB alt stack a direct leak: nothing retains the
pointer once sigaltstack() owns it. Map it instead of allocating it, which
also keeps the handler's own stack out of the heap whose corruption it may
be reporting.
The leak was the tell for the worse bug. ASan already has an alt stack
installed by the time signal_handlers() runs, 32 kB of it, and the old size
floor rejected it and installed ours over the top, moving ASan's handlers
onto our mapping. Defer to any alt stack already installed regardless of
size: whoever put it there sized it for their own handler, and ours needs
about 10 kB of the 32 that the smallest of them offers.
Test 181 pins both directions through an LD_PRELOAD observer that compares
the mapping the process ends up with against the one it was handed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Refresh mergeability after master moved
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Installed dev headers do not compile standalone
Three headers in DevIncludes_DATA fail on their own from $(includedir)/httrack
once HTS_INTERNAL_BYTECODE is defined: htswrap.h includes coucal.h, which is
never installed and which the file does not use; htsmodules.h uses HTSEXT_API
without htsglobal.h; htsdefines.h uses size_t without <stddef.h>.
205_install-headers.test installs the list into a temp dir and compiles a
one-line TU per header in both preprocessor states.
Closes#943
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Split CC so the test runs on the linux-i386 leg
That job configures with CC="gcc -m32", and looking the whole string up as one
executable made the test exit 77. A skip reads as success, so the one leg with a
different data model was checking nothing. Same for CC="ccache gcc".
The compiler now also has to compile a bare <stdio.h> first, or a box without
gcc-multilib blames every installed header instead of reporting its own toolchain.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Pass the build's CPPFLAGS to the installed-header compile
htsbasenet.h includes <openssl/ssl.h> under HTS_USEOPENSSL, so the test needs
OpenSSL's include path like any consumer does. Linux resolves it from
/usr/include, which is why only the macOS leg failed: brew's openssl@3 is
keg-only and reachable only through the CPPFLAGS given to configure.
Reproduced by hiding /usr/include/openssl behind a -nostdinc mirror and exposing
it from a private prefix, which reds htsnet.h and htsopt.h exactly as CI did.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Tooltip JavaScript breaks on any translation containing an apostrophe
The onMouseOver handlers interpolate through ${html:...}, which escapes
for HTML text. The browser decodes an attribute value before compiling
it as JavaScript, so ' arrives as a bare quote and closes the string
literal; a translation carrying a double quote ends the attribute
outright. Add a js: template mode escaping for both layers and move all
304 handler sites onto it.
Closes#864
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Close the step2.html and title-attribute gaps in the js: escaping test
The static check could not match ${html:html:...}, the very shape it was
added to catch, and the runtime probe only fetched option1.html: reverting
all eight step2.html handlers left the test green. Nothing asserted that a
title= attribute stays on html: outside option1.html either.
Walk every server template instead, pairing each ${...} with the attribute
holding it, and fetch step2.html as well. Also feed the fixture a newline
and a tab, which had no coverage.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Emit the backslash as \x5c so a DBCS trail byte cannot swallow it
shift-jis, BIG5 and gb2312 all accept 0x5c as a trail byte, and those are
the declared charsets of three shipped lang files, so the browser really
does run a DBCS decoder over the page. An orphan lead byte followed by the
emitted \\ pair decodes as one glyph plus a stray backslash, which escapes
whatever comes next. Emitting every escape as a \xNN group leaves the
decoder nothing to pair with: the only bytes cat_js_escaped adds are '\',
'x' and hex digits.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Nine lang.def strings are untranslated in 26 of the 30 language files
LANG_F15b, LANG_I6c, LANG_I23c/d/e, LANG_I35c and LANG_I43c/d/e were
translated only in English, Francais, Dansk and Portugues-Brasil, so the
Windows GUI's option dialogs fell back to English everywhere else. Fill
them in for the other 26 files, each encoded in that file's declared
LANGUAGE_CHARSET and inserted where English.txt orders it.
Svenska.txt declares ISO-8859-2 but holds Latin-1, and Chinese-BIG5.txt
needs cp950 rather than strict big5. Romanian.txt and Slovenian.txt
declare ISO-8859-1, which cannot spell either language, so both stay
unaccented ASCII the way #862 left them.
Test 62 waived these nine msgids by name and fails once a waiver is no
longer needed, so drop them from that list: it reports 231 untranslated
msgids against master.
Closes#863
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Romanian: restore the two circumflexes ISO-8859-1 can spell
Only the-breve, s-comma and t-comma are unrepresentable in the declared
charset; a- and i-circumflex are 0xE2 and 0xEE, and the rest of the file
already spells them that way.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Icon=httrack never resolves through the share/pixmaps fallback
share/pixmaps only ever held the sized names, so the bare-basename lookup a
desktop falls back on when no icon theme is installed never matched. Install
httrack.xpm (a 32x32 copy) alongside them rather than renaming one: Debian
globs the whole directory and a rename would strand the old names on upgrade.
Closes#932
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Pin the pixmaps fallback to the icon, not to any 32x32 image
The geometry check passed on a solid-red square, so compare the installed
fallback against the installed 32x32 copy, declaration line aside. Also drop
the banned pipe into grep -q while here.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* The suite watchdog may kill its own taskkill, and an empty category reads as a failing test
ci_suite_heartbeat runs in a subshell of the process it targets, and kill_tree
on Windows is taskkill /F /T /PID, so taskkill is a grandchild of its own
target. A leaves-first /T would reap the watchdog before reaching the root: the
annotation is printed, nothing dies, and the step runs on to the 45-minute
cancel with no log. Signal the target directly first, then through the tree, so
the kill no longer depends on taskkill outliving its own ancestors.
The driver's category globs also expanded to themselves when they matched
nothing, handing test-timeout.sh a literal pattern that exits 127 and lands in
the tally as a failing test named after the glob. Categories now carry a label
and expand under nullglob; an empty one names itself and stops the suite before
any test runs.
Closes#952Closes#953
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Make the empty-category gate cover the one category that names a single test
runnable:00_runnable.test carries no metacharacter, so nullglob cannot empty it
and the gate never fired for it: the name reached test-timeout.sh unexpanded and
was counted as a test failing 127, exactly what the gate exists to prevent. Give
it a pattern like every other category, and empty it in the driver test.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
tools/Info.plist is the one that matters: configure generates it from
Info.plist.in, so a stray "git add -A" would commit a snapshot that
shadows it and pins CFBundleShortVersionString to whatever version built
it (#884, one directory over).
Closes#905
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
#897 pinned the metainfo version against htsglobal.h, but nothing checked
the file was valid AppStream: a broken one passes the suite and a software
centre then drops WebHTTrack without saying so.
The lint job gains an appstreamcli step beside the .vcxproj well-formedness
pass, and 210_appstream-metainfo.test runs the same check under make check,
skipping when appstreamcli is absent. --no-net keeps the gate off the
screenshot host's reachability; errors and warnings fail, hints do not.
The test validates a copy with <id> stripped as a positive control, so a
validator that never fails cannot make it vacuous.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
windows-build.yml was the only workflow without a concurrency group, so
every push left the runs before it going. Match ci.yml and codeql.yml.
Closes#942
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Derive the bundle's minimum macOS from its payload
tools/Info.plist.in hardcoded LSMinimumSystemVersion 11.0, but nothing
built the payload for 11.0: with no MACOSX_DEPLOYMENT_TARGET set, clang
on the macos-15 runner targets the host, and all nine Mach-Os in the
released DMG report minos 15.0. Finder believes the plist and launches
on macOS 11 through 14, then dyld refuses the binaries.
macos-app.sh now takes the highest minos in the assembled bundle and
writes it to LSMinimumSystemVersion before the ad-hoc signature seals
Info.plist, so the two cannot drift again. The floor has to be the
maximum rather than the main executable's: five of those nine are
Homebrew bottles whose minos was fixed when the bottle was built, so a
deployment target on our own compile would not move them.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Assert the declared minimum macOS instead of rewriting it
Review pushed back on deriving LSMinimumSystemVersion from the payload:
it makes the shipped minimum follow whatever runner image built the
bundle, so bumping a CI label would move HTTrack's advertised floor with
nobody deciding it. Declare 15.0, which is what the binaries need today,
and fail the build when the payload disagrees.
macho_floor also swallowed an otool failure: the call sat non-final in a
pipeline, so a Mach-O otool could not map contributed nothing and the
floor came from the rest, which is the same wrong-floor bug this fixes.
deps() already guards its otool the same way.
Dropping the plist rewrite drops plist_set with it, and with it the
awk-on-XML edge cases and the read-back check that was only meaningful
while the derived value differed from the template's.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Trim the comments the review flagged
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <xroche@gmail.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* The CI watchdog deadlines stretch under the starvation they exist to catch
Every deadline in the test harness counts poll iterations and assumes each
costs exactly the tick it asked for. Under the CPU and fork starvation that
wedges a runner, an iteration costs several times that, so the budget arrives
late in wall-clock terms or never: the 45-minute step timeout cancels first,
and a cancelled step keeps neither its log nor its artifacts. That is why the
wedged Windows jobs in #795 name no test.
Measure $SECONDS instead, in test-timeout.sh and in run_with_timeout,
wait_bounded and reap_bounded, which restores the wall-clock contract
run_with_timeout's comment already described.
The guards are checked against a reverted copy via a slow-sleep shim: with
each poll stretched fourfold, the counting code takes 44s to honour a 1s
budget and 16s a 3s one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Compare the deadlines strictly, and make the guards kill their mutants
$SECONDS is floored, so a reading equal to the budget can be a fraction
under it: -ge fired up to a second early, which kills healthy work. Compare
strictly instead, at the cost of at most a second late.
The 58_watchdog guard also passed on a watchdog that killed everything on
sight, because one stretched reap poll lifted the elapsed time over its lower
bound, and its control raced (the child exited before the first poll). Give
the control a command that outlives several polls, and widen the timing bound
to sit a full poll quantum either side of 8s healthy and 16s counted.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
tools/macos-release.sh signs every Mach-O with the hardened runtime, notarizes and staples the app, packs a DMG and notarizes that too. A new macos-release workflow runs it on a tag or on demand, with a Developer ID imported into a throwaway keychain from repo secrets. Signing is per-Mach-O rather than --deep, which reaches nested code but never applies the runtime, and notarytool's status line is read rather than its exit code.
The bundle stub is now a Mach-O (tools/httrack-launcher.c) rather than a shell script, which Apple treats as a resource and TCC cannot attribute a prompt to. The Mach-O walk and the Info.plist reader are shared with macos-app.sh via tools/macos-bundle.sh, and ci.yml's macos-app job runs the same script ad-hoc on every PR.
Closes#901
* The Windows suite driver is 140 lines of shell inlined in YAML
Move the "Run the engine test suite" step body to tests/ci-windows-suite.sh,
where shellcheck and shfmt reach it and it can be run by hand. The two
deliberate word-splits in the skip-set compare needed a directive; nothing
else changed, verified by diffing the shfmt-normalized old body against the
new file.
ci_annotate and ci_suite_heartbeat move with it, out of the test library
every test sources on every platform. 171_watchdog-heartbeat.test sources
the driver, which returns early unless run directly.
Closes#948
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Test the driver on this platform, and ask the shell whether it was sourced
The guard read "${BASH_SOURCE[0]}" = "$0", which is only false by accident of
what the caller put in $0: bash -c '. "$0"' <driver> makes the two equal and
falls through into the suite. Ask the shell instead.
172_ci-windows-driver.test pins that case, the bindir contract and the loop's
accounting against a stub bindir, none of which a Windows-only leg proves
before merge.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Silence two shellcheck findings in the new driver test
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Make the driver test survive distcheck's read-only srcdir
cp carries the source's mode over, so under distcheck the neutered testlib
copy came out read-only and the append failed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* Bundle the non-system dylibs, or a downloaded HTTrack.app cannot launch
src/Makefile.am links $(OPENSSL_LIBS) into libhttrack, which on a build
machine resolves to Homebrew. The bundle check only rejected the staging
prefix, so /opt/homebrew paths sailed through: fine for someone who ran
brew install, fatal for anyone who mounts a DMG.
macos-app.sh now copies the transitive closure of non-system dylibs into
Contents/Frameworks, rewrites the load commands to @rpath, and adds a
depth-correct @loader_path rpath to every Mach-O. install_name_tool
invalidates the signature and arm64 refuses to run an unsigned binary, so
each rewritten file is re-signed ad-hoc.
The check is now that every load command resolves to /usr/lib,
/System/Library, or a file present in Frameworks, and it fails if it
scanned no Mach-O at all. Because a static check cannot fail while the
loader is quietly falling back to Homebrew, CI also hides Homebrew's
openssl@3 and re-runs the smoke against the moved bundle.
Part of #901.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Give the linker header room for the @rpath rewrite
install_name_tool refuses when the new load commands do not fit the
existing header: "changing install names or rpaths can't be redone".
The bundle build now passes -Wl,-headerpad_max_install_names, and the
script says so when the rewrite fails rather than surfacing the
toolchain's own message.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Fix two bundling bugs the review found, and three checks that could not fail
machos() ended with a while loop, so its status was the last iteration's
file(1) test. The bundle always holds shell scripts, and find walks APFS in
directory-hash order, so whenever a non-Mach-O came last the script died
under set -e with no diagnostic and a half-populated Frameworks. Green CI
here was luck, not evidence.
Two dependencies sharing a basename were both copied in the same pass, the
second clobbering the first, and relink() then pointed both references at
the one survivor. Frameworks is flat and cannot express the difference, so
this now fails with both paths named.
The checks that could not fail:
- machos() prefiltered on mode and suffix, and the same list drove the copy,
the relink and the audit, so a non-executable Mach-O was unbundled AND
unchecked. It now enumerates every file.
- deps() piped otool into awk, so a failing otool yielded no dependencies and
the binary was recorded as clean. Its status is now checked. Reading the
indented lines rather than tail -n +2 also stops a universal binary's
per-slice headers from parsing as dependencies.
- The CI probe hid only the openssl@3 opt symlink, which a source-built
bottle does not reference, and asserted nothing about the bundle first, so
an empty Frameworks passed it. It now asserts libssl is bundled, then hides
the kegs of every library in the closure.
- Signing verified the bundle after codesign --force --deep had already
re-signed it, repairing the damage it was meant to catch. Per-file
signatures are now verified first.
-Wl,-headerpad_max_install_names moves from the CI job to configure.ac: it
is a precondition of the macos-app target, not a CI preference, and anyone
following the standalone recipe in Makefile.am:22 hit a rebuild-only
failure that ld64's default header slack made intermittent.
The toolchain gate is now unconditional, so a non-Darwin host refuses to
build rather than emitting a bundle nothing verified.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* One library reached by two paths is not a basename collision
The bundle depends on both /opt/homebrew/opt/openssl@3/lib/libcrypto.3.dylib
and the Cellar path behind that symlink, so comparing the dependency strings
called Homebrew's own layout a collision and refused to build.
Compare the paths with the directory resolved instead. Two names for one
file now dedupe; two different files under one name still fail, which is the
case that would silently clobber.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The Windows job wedges every few dozen runs: the test step never completes, the runner is lost, and the log blob and the `if: always()` artifact uploads go with it. So the wedging test has never been named.
A watchdog now runs beside the suite and ends the step before the runner dies, because a step that fails on its own terms keeps its log and still runs the uploads. It never triggers on elapsed time, since a healthy test and a wedged one look identical by the clock. It watches the progress log instead, where every outcome writes a line, the per-test timeout included: 900 seconds without one means that timeout did not fire, which is the wedge. On the way it names the test in flight as annotations, which is live progress rather than evidence.
That distinction is measured, not assumed. A throwaway workflow emitted notices and then died four ways: a step cancelled by its own timeout keeps its log and every annotation, on Linux and on Windows, from a background subshell as well as the foreground; a runner killed mid-step keeps neither, dropping notices that had been on the wire for twenty seconds.
Refs #795 rather than closing it: this ends the silence, it does not fix the leak behind it. Leftover coverage gap in #949.
Bump src/coucal from 5d2a633 to 0d36322, taking the two commits #945
deliberately skipped.
#31 replaced the log-level #if 0 cascade with COUCAL_LOG_LEVEL, but
defined all five level emitters unconditionally, so the unused ones
warned under -Wall for any consumer lacking coucal's own
-Wno-unused-function. That was the reason #945 pinned behind head.
xroche/coucal#33 tags them with an unused attribute instead and drops
the exemption from coucal's Makefile, so the warnings are gone at the
source rather than masked, and a new upstream CI leg compiles coucal.c
with consumer flags at every verbosity.
The default verbosity is still info, so the compiled-out debug and
trace that #941 measured stay compiled out: the probe from that PR
still reports 0 print handler calls per 20000 inserts. Our build is
clean of compiler diagnostics under gcc and clang. MSVC is unaffected
either way, since it builds at /W3 and the equivalent C4505 is level 4.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* proxytrack: bound the .arc reader and stop the writer trusting a size it has no bytes for
The reader hands back an element carrying a declared size with adr == NULL when
it could not fetch the body, and the .arc writer took the size at face value:
fwrite(NULL, 1, size) faulted inside libc (#931). Both writers now take the
body from a helper that answers 0 when there is nothing to write, and the
record's own length, the Content-length header and the md5 follow it, so what
is written declares what it holds.
A new fuzz-arc harness drives the reader the way --convert does (#929), and
found the rest in seven executions: PT_Delete never freed the indexes it owns,
PT_Index_Delete__Arc never freed its hashtable, and a record could declare a
two-gigabyte body that the reader allocated before the short read failed.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* proxytrack: commit the merged index slot only once the array holds it
PT_IndexMerge() counted the new slot before growing the array, and assigned
realloc's result straight onto indexes->index: a failed allocation dropped the
array it already had and left index_size counting an entry that was never
stored. Harmless while nothing walked the array; PT_Delete() now does. The
array holds pointers, so size it as such rather than as whole PT_Index structs.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* proxytrack: trim the comments added by the .arc hardening
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* proxytrack: initialise the .ndx mutex, keep unsizeable archives loading, and give the bound a positive control
PT_LoadCache__Old() never called MutexInit(), which was survivable while
PT_Index_Delete__Old() was unreachable; PT_Delete() now runs it on every exit.
The readers lock the same handle, so on Windows the .ndx path was serving
unsynchronised.
A .arc past LONG_MAX cannot be sized by a 32-bit ftell, and refusing the whole
archive lost the records that used to load; the bound just stops constraining.
Test 164 asserted only refusals, so a reader that refused every body passed it,
as did an off-by-one on the last record. It now converts an archive whose body
ends on the last byte of the file, and checks the zip writer's own size header.
The truncated.arc seed ran out of file before reaching the bound it is named
for.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <xroche@gmail.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Bump src/coucal from a0a9e49 to 5d2a633, clearing the one warning #941
knowingly landed: 7a8198d const-qualified the_empty_string but still
returned it through a plain cast to coucal_key (void*), which httrack
compiles with -Wcast-qual. Upstream keeps the const and routes the
singleton through uintptr_t, so it stays in rodata, and adds
-Wcast-qual to its own mandatory flags so the class cannot return.
Pinned at 5d2a633 rather than upstream head. The next commit (#31)
replaces the log-level #if cascade with COUCAL_LOG_LEVEL and defines
all five level functions unconditionally, which warns twice under
-Wall for any consumer without coucal's own -Wno-unused-function.
Filed as xroche/coucal#32; head follows once that is resolved.
The default level is still info, so the compiled-out debug and trace
that #941 measured stay compiled out: the same probe reports 0 print
handler calls per 20000 inserts at both 5d2a633 and head.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Bump src/coucal from 93ec411 to a0a9e49, five upstream fixes.
The one that reaches the engine is the logging change. coucal_trace()
in coucal_add_item_() passes coucal_print_key() as an argument, and the
compiled-out macro expanded to an ordinary variadic call, so the
argument still ran. htshash.c installs key_adrfil_debug_print on
hash->adrfil and hash->former_adrfil, where it snprintf()s the full URL
into a scratch buffer, so every insert into the dedup tables paid for a
URL format that was then discarded. Measured on a probe mirroring
htshash.c's setup: 13262 handler calls per 20000 inserts, now 0, with
the table statistics unchanged.
The other four fixes harden coucal without reaching our call sites: the
custom key free at destruction (our dup handler is an identity pointer
copy and the free handler is empty), the mid-walk delete enumeration
skip (all four enum loops in htsback.c are read-only), the pool-aliased
key use-after-free (no call site passes an item name back as a key),
and the coucal_new() shift width (every call site passes 0).
No ABI change: coucal.h is not installed, struct_coucal is opaque,
struct_coucal_enum is byte-identical, and the 36 coucal symbols
libhttrack exports are unchanged.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
#933 and #934 merged minutes apart and both landed a test numbered 160, so the icon-theme one becomes 163. Only `tests-list.mk` referred to it, and the Windows job picks tests by topic word rather than number, so its coverage is unchanged.
The cache-hdrbounds test also sent stderr to `/dev/null` to hide an expected warning, which threw away the self-test's own diagnostics with it: a failure printed `cache-hdrbounds: FAIL` and nothing about why. stderr now goes to a file both failure paths report.
Four places built or consumed the cache key, each assuming a different maximum URL length. A URL long enough to fill a `lien_back` field is legal and reachable off the wire, so one of them aborted the crawl where it should have missed the cache, and the index load read entry names into a buffer minizip can fill without a terminator.
All four now size off one bound, and the key is built all-or-nothing with the existing `slcatprintfbuff`: too long to store drops the entry with a warning, too long to look up is a miss. Clipping is never right here, because a clipped key is a valid key for some other URL, and that is a cache hit on the wrong content. The new `cache-urlbounds` self-test (`tests/162`) stores at the cap and pins that neither a twin differing only in its last byte nor a decoy sitting on a clip point can be served in its place.
Closes#935Closes#936
* Bound the cache header block instead of trusting the field caps
ZIP_FIELD_STRING and its integer siblings sprintf'd into a fixed 8192-byte
block with no bound, in both the engine cache writer (cache_add) and
ProxyTrack's new.zip writer. The values are remote-controlled -- ETag,
Location, Content-Disposition, the URL itself -- and in cache_add the caps
they are declared with sum past 8192, so the writer could overrun its own
stack buffer.
Route every field through slcatprintfbuff(), a new all-or-nothing bounded
append: a field that does not fit is dropped whole, since a clipped one reads
back as a valid shorter value. X-Save moves ahead of X-Addr/X-Fil so the one
field the reader consumes is not the first casualty of a full block.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Merge origin/master into fix-841-zip-field-bounds
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Restore the bounded ZIP_FIELD_STRING lost in the merge
The merge commit picked up a mutation-testing revert of this macro from the
shared worktree, putting the unbounded sprintf back.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Make the header-bounds test kill the mutants it was walking past
A test audit built seven mutant writers against the new self-test; five
passed. Restoring the legacy field order passed while silently dropping
X-Save, so the reorder this PR relies on was asserted by nothing. An
early-return writer passed vacuously, the check landing on the control
entry because nothing pinned which entry it read. A truncated Location
passed because any non-empty prefix was accepted, and the size bound was
a literal 8192 decoupled from the buffer it was meant to track.
The block is now identified by a field only the maxed entry carries, the
bound comes from a shared CACHE_HEADERS_SIZE, X-Save must survive, and a
still-present X-Fil reports that the entry stopped filling the block
rather than passing quietly. The wrapper consults httrack's exit status,
so an abort after the verdict is no longer a pass.
Adds an .arc to .zip round trip for ProxyTrack's writer, which had no
runtime coverage: every --convert in the suite writes .arc.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Add tools/HTTrack.icns from the brand master, declare it in Info.plist.in,
copy it into the bundle, and check the plist and the payload agree.
Closes#900
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* The application icon is still the pre-brand bitmap set
Replace the 16/32/48 PNGs and the .xpm fallbacks with the HT monogram
generated from the same Jost* master as the masthead wordmark, and extend
the hicolor theme with 64, 128, 256 and a scalable SVG.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Assert the icons on what install and dist emit, not on Makefile.am text
Six mutants survived the first version: an emptied EXTRA_DIST, an empty *dir
variable (automake's install rule exits 0 when it is), a _DATA glob with the
wrong extension, a _DATA entry for a directory the tree does not have, a
missing apps context subdirectory, and a size dropped from install entirely.
All six now fail, and the PNGs are indexed rather than RGBA.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* proxytrack cannot re-read the .arc it writes
The version block's declared length counted the blank line closing it, so
the reader consumed the first record's separator and every entry was
skipped: a second --convert over proxytrack's own output loaded nothing.
The bytes on disk are unchanged; only the declared length shrinks by one,
which an older proxytrack reads too. The reader now stops on the last of
the newlines closing the version block, so archives already written the
old way still load.
Closes#834
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Bound the version block newline scan and keep rejecting truncated archives
Review of the first commit found two regressions of its own: the newline
run was scanned to its end, so an archive padded with a gigabyte of them
cost a gigabyte of reads where master stopped after two, and tolerating
EOF there turned a length running past the end of the file into a silent
empty load. At most two newlines are read now, and a version block that
does not end on one is rejected as before.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Pin the version block length against a compensating extra newline
A writer that emits the blank line and still counts it round-trips, so
every assertion passed while the length stayed a byte too long. The
declared block must not end on a blank line either.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* The tagline no longer sits against the wordmark's baseline
The old GIF was cropped from the cap tops to the baseline, so its box edges were
the letters. The SVG's box is the true ink box, which in Jost also holds the k's
ascender above the caps and the round letters' overshoot below the baseline, and
that shows up as a band of field colour between the wordmark and the tagline.
Neither band can be cropped out of the artwork without cutting ink, so the
masthead takes them back optically. Measured against the old rendering, the
baseline-to-tagline distance is identical and the cap tops land within a third
of a pixel.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Cut the comments back to one line each
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Say what the margins actually do
The comment claimed both bands were trimmed. The lower one is, in full; the
upper is trimmed only by what the old bitmap did not already carry, which is why
1.7px is not the 3.56px the artwork measures. Derivation recorded beside the
generators.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <xroche@gmail.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* ProxyTrack dumps a coucal hashtable stats line on every WebDAV request
coucal logs a per-table statistics summary when a table is deleted, and
with no handler installed it prints that line itself, prefixed with the
table's address. ProxyTrack builds and drops one table per WebDAV
enumeration, so every unauthenticated PROPFIND put one on whatever the
service redirects its output to, heap pointer included.
httrack and htsserver were already covered by hts_init(), which installs
a global coucal handler that drops info-level messages unless HTS_LOG is
set (#416). ProxyTrack does not link libhttrack: it compiles coucal
itself and never calls hts_init(), so it was the last binary on coucal's
built-in sink. It now installs its own handler. Critical and warning go
through proxytrack's log, everything below is dropped unless HTS_LOG is
set, and the summaries then come back without the address.
Unnaming the tables would not have fixed it. coucal_delete() logs the
summary for every table, named or not; the name only decorates the
message.
Closes#918
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Whitelist proxytrack's request-time output instead of naming two absent strings
The absence check pinned two literals, so a reworded leak or one whose
pointer was not at column 0 walked through it. Assert instead that every
line the process writes once serving is an access-log line, which is the
property, and pin the routed form under HTS_LOG whole so the address
cannot creep back between the severity and the message.
Also say why the handler passes "debug" rather than the DEBUG macro: the
macro is NULL outside a debug build, which would quietly make HTS_LOG a
no-op.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Order the whitelist against a request's teardown, not just its access log
proxytrack writes the access line before send(), so waiting on it ordered
nothing that a teardown or keep-alive path writes afterwards: a leak
delayed 400ms past the response survived the check. Send both PROPFINDs
over one connection, since the keep-alive loop does not read the second
request until the first one's teardown has run, and read the capture only
once proxytrack is gone and the pty drainer has marked it complete.
Draining alone was not enough. SIGTERM cuts the work short rather than
truncating a buffer, so a still-pending write is never made at all and
there is nothing left to flush; the ordering is what catches it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
A FIFO passes "test -x", so the #920 check ran it: bash gets EACCES from
execve, falls back to reading the file for a shebang, and blocks in open()
with no writer. AS_EXECUTABLE_P is autoconf's own "test -f && test -x", and
AC_PATH_PROGS on the line above already applied it to the PATH search, so the
override path was simply using the weaker predicate.
A regular executable that never returns stays uncovered: no portable timeout
is worth it, and CC= pointing at the same wrapper hangs stock AC_PROG_CC too.
Test 151 gains the FIFO case, and its run() is capped so a regression fails
instead of wedging "make check" with no log.
Closes#922
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* http_xfread1's reserve-only read mode has no caller
The `bufl == -2` branch of `http_xfread1()` allocates the line buffer and
returns without reading. Nothing has ever called it: no call site in the
tree passes -2, and scanning all 2649 revisions in this repository for
`xfread1(` call sites turns up 24 distinct lines, none of them -2. The
branch arrived with the 3.20.2 import commented "force reserve", so it was
probably meant for a preallocate-then-fill pattern that never landed.
Naming it `HTS_XFREAD_RESERVE` in #919 made it read as a supported mode.
No external caller is possible either: `htslib.h` is not installed and the
symbol is hidden, so this is not an API change.
Equivalence checked against the object code. `htsback.o` and
`htsselftest.o` disassemble identically; in `htslib.o` every function
except `http_xfread1` differs only in the `__LINE__` values `htssafe.h`
bakes in, shifted by the seven deleted lines. No new test: nothing changes
for any input a caller can produce, and `01_engine-xfread` plus the chunked
tests still pass.
Closes#923
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Say that any non-positive bufl is line mode, not just the two named
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0199wAkSVZNBNp51mpRkxMvv
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* The masthead wordmark is a 400x34 GIF that blurs on any hi-DPI screen
Replaces it with an outlined SVG across the 38 documentation and WebHTTrack
pages that carry it. The original was set in Futura, so the lockup was refitted
in Jost*, the closest free Futura revival, taking weight from the measured stem
thickness, size from the cap heights and tracking by least squares against the
glyph positions in the old bitmap.
tests/82 now asserts that every image a GUI page names is actually served.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Point the shared chrome generator at the new wordmark
The masthead of the 13 generated pages comes from tools/doc-chrome.py, so
editing the pages alone left the generator disagreeing with its own output and
--check red.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <xroche@gmail.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* proxytrack: drop the leftover debug traces on the WebDAV path
Two fprintf(stderr) calls in the PROPFIND path shipped by accident: one dumped
the client-supplied request body, the other the whole generated multistatus
response. Both ran on every PROPFIND with no authentication in front of them,
so any client could write bytes of its choosing into the operator's stderr.
The body is never parsed and the response is derivable from the index, so
neither trace has diagnostic value worth keeping behind a debug level.
Closes#911
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: make the #911 leak check see stdout and the property, not two literals
The absence checks pinned 'DEBUG: DAV-DATA' and '^RESPONSE:', which four
re-added variants walk straight past: a trace without the hyphen, one with no
marker at all, one prefixed so the '^' misses, and one on stdout. The stdout
case is the worst of them: redirected to a file, proxytrack's stdout is fully
buffered and SIGTERM never flushes it, so the leak never reached the log the
test reads.
Give proxytrack a pty instead of a file, so libc line-buffers its output on
Linux and macOS alike, and assert the property: a PROPFIND carries a canary the
index cannot produce, and neither the canary nor a distinctive string from the
generated response may appear in what proxytrack wrote. The two literals stay as
names for the specific regression. The liveness guard now requires a PROPFIND
answered 207, since a depth-rejected one is logged 403 by the shared reply path
without ever reaching the deleted code.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Bound the ProxyTrack DAV item buffer against an amplified PROPFIND path
proxytrack_add_DAV_Item() reserved a fixed 1024 bytes and then sprintf'd into
it unbounded. The request path lands in the response twice, once as the href
and once as the displayname, and escapexml() turns each '&' into '&', so
an unauthenticated PROPFIND of roughly 900 ampersands writes about 9000 bytes
off the end of the heap block. No cache entry and no Depth: 1 are needed.
Replace the hand-sized reserve with StringSprintf(), which measures the
formatted output and grows the String to fit, and convert the sibling sprintf
sites in the same file so no unbounded write into a String is left to
re-audit. Sizing beats clipping here: the String already owns a growable
buffer, so nothing has to be dropped.
Closes#836
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Bound StringSprintf's pre-C99 retry, and trim the review findings
A genuine vsnprintf conversion error returns -1 just as pre-C99 msvcrt does
for a short buffer, so the doubling search had no way to tell them apart and
grew until realloc aborted. Unreachable from these format strings, which use
only %s and %d, but the helper lives in a shared header and will get more
callers. Cap the search and empty the String past it.
Also: the count assertion piped into wc under pipefail, so a zero count killed
the test through set -e before its diagnostic could print.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Lift the NO_WEBDAV conditional out of a macro argument list
A preprocessor directive inside a macro invocation's arguments is undefined:
it was fine while this was a plain sprintf() call, and MSVC rejected it as
soon as it became StringSprintf(). GCC accepts it, so only the Windows leg
caught it. Compute the DAV header fragment first and pass it as an argument.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Cover StringSprintf's exact-fill case and the WebDAV enumeration branch
StringSprintf_ writes the terminator at buffer[ret], so widening its
`ret < capacity` guard by one byte is a heap overflow that only fires when the
formatted output exactly fills the capacity. No crawl test lands on a
capacity boundary, so the mutant survived the suite. The new `strsprintf`
self-test sweeps lengths around 256, 512, 1024 and 2048 with the String's
capacity pinned to each, plus a growing and shrinking sweep on one reused
String, and checks the length, the bytes and the terminator every time.
Test 147 only ever sent Depth: 0, leaving the enumeration branch the same PR
rewrote with no coverage at all. Its fixture gains a child directory, and a
Depth: 1 listing pins the item URLs, including the trailing '/' that
StringPopRight takes back off a directory name.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Make the String failure paths safe without assert
StringSprintf empties the String when it gives up, but that contract was
only visible in the implementation, and the WebDAV enumeration in
proxytrack pops the trailing '/' straight after it. State it at the
declaration, no-op StringPopRight on an empty String, and skip an
enumerated item the formatter could not name.
StringRoomTotal reported a failed realloc through STRING_ASSERT alone.
The MSVC Release configuration defines NDEBUG, so that check is already
gone from the shipped Windows builds, leaving a NULL buffer under a
capacity bumped before the allocation was known to succeed. Assign both
only on success, and terminate through StringOom_.
Closes#915
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Renumber the String OOM test to 152
151 is taken by the unmerged tests/151_bash-shell-validate.test (PR #920).
The filenames differ, so git would have carried both onto master rather than
conflicting.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Declare the new WebDAV test in the Windows skip set
It skips on Windows for the same reason as its two neighbours, MSYS
cannot reap a background listener (#595), and the ratchet fails a skip
it was not told about.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0199wAkSVZNBNp51mpRkxMvv
Signed-off-by: Xavier Roche <roche@httrack.com>
* Keep the out-of-memory action overridable
STRING_REALLOC and STRING_FREE are #ifndef hooks, and STRING_ASSERT was one
too; replacing it with a hard-wired call took a hook away from downstreams of
this installed header. Route the failure through STRING_OOM instead, with the
print-and-abort default unchanged.
Also flush stderr before aborting: the Windows CRT buffers a redirected
stderr and abort() flushes nothing, which would drop the message the test
matches on.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Inject the allocation failure instead of asking for a huge one
The engine self-test pinned a String's capacity so the next doubling asked
for SIZE_MAX/2, on the assumption that no allocator would serve it. Six CI
legs disagreed: i386 has a 3G user space, and the 64-bit runners handed the
request out too, so the test reported "NOT aborted" everywhere but here.
Green that depends on how much memory the machine feels like giving is not a
test.
Drive the path from a standalone helper instead, which defines STRING_REALLOC
to a stub returning NULL before including htsstrings.h. Four cases: growth
with the stub allocating for real, the failure reaching the handler with the
size it asked for, a live buffer surviving a failed realloc, and the shipped
handler printing and aborting. Only the automake build produces the helper,
so the test declares its Windows skip.
The self-test had no portable way to force the failure, so it goes rather
than staying as a handler nobody can rely on.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Assert the bytes and the requested size, not just the bookkeeping
An under-allocation survived the helper: shortening the realloc by one byte
while still recording the full capacity left all four cases green, because
only the growth case allocated anything and it checked the capacity number
rather than the memory behind it. Fill the announced capacity to its last
byte and read it back, which the sanitizer legs turn into a hard failure.
The failure cases pinned the initial capacity by asserting 16, so bumping
that policy would have failed a correct tree. Compare the size handed to the
handler against the size the stub was actually asked for instead, which also
catches the under-allocation on legs with no sanitizer.
Drive StringSprintf_ and StringBuffN_ too, the other two places the header
expands STRING_OOM.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* configure accepts a BASH_SHELL that is not a usable bash
`./configure BASH_SHELL=/bin/sh` was accepted without a word. `AC_PATH_PROGS`
takes any absolute value verbatim, so the macOS problem #895 fixed (a bash in
POSIX sh-mode driving `make deb` and the test harness) came back, surfacing
much later as a `146_bash-shell.test` failure instead of a configure error. A
relative value never reached the Makefiles at all: it was dropped for whatever
the PATH search turned up.
configure now checks the value it resolved. The shell must be executable,
report a `BASH_VERSION`, and not carry `posix` in `SHELLOPTS`, the
discriminator `146_bash-shell.test` already uses, since an sh-mode bash reports
a version too. A relative or whitespace-carrying override is refused before the
search runs; whitespace would otherwise survive into `$(BASH_SHELL)`, which
nothing in the Makefiles quotes.
Only an explicit override is fatal. When the search itself finds nothing
usable, configure warns and carries on, so a box without bash still builds; it
just cannot run `make check` or `make deb`.
Closes#908
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* An environment in POSIX mode must not be blamed on the bash path
POSIXLY_CORRECT, or an exported SHELLOPTS, puts every bash into POSIX
sh-mode, so `./configure BASH_SHELL=/bin/bash` failed with advice to pass a
path that cannot exist, and a plain configure warned that no usable bash was
found on a box that has one. The probe now runs a second time under `env -u
POSIXLY_CORRECT -u SHELLOPTS`; if the shell is fine once they are cleared, the
message names them and says how to clear them for make as well, since it
inherits the environment. An override stays fatal, the search still only warns.
The bash-ness probe read `BASH_VERSION`, an ordinary variable any shell echoes
back, so `BASH_VERSION=9.9 ./configure BASH_SHELL=/bin/dash` was accepted and
dash landed in `TEST_LOG_COMPILER`. It reads `${BASH_VERSINFO[0]}` instead,
which no environment can fake.
The path guard covered whitespace alone while claiming to cover what make and
the recipe shell split on, so a real bash under a directory named with `;` or
`$` or `#` still reached the Makefile. It now rejects that whole class. Quoting
`$(BASH_SHELL)` at its three uses was the alternative, but make cuts the value
at a `#` and expands a `$` before any shell sees it, so quoting would cover
less than the guard.
The suite now also pins the branch the fatal/warn split rests on: no override,
an unusable bash first in PATH, configure exits 0 with a warning.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Nothing pinned which cause of POSIX mode gets blamed
A shell can be in POSIX mode because it was invoked as sh or because the
environment forces it, and only the second probe can say which. Nothing held
that down, so a version deciding from the environment alone, without
re-probing, passed every case in the suite while telling the user to clear a
variable that would not have helped. Test 151 now runs a bash symlinked as sh
with POSIXLY_CORRECT=1 set as well, and requires the message to name the path.
The comment records why that branch cannot simply read POSIXLY_CORRECT:
autoconf runs "set -o posix" on configure's own shell, so it is set there no
matter what the user's environment holds.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* A chunked response carrying trailers is discarded as "Invalid chunk"
The chunk automaton expected the line after the terminating zero-length
chunk to be empty. RFC 9112 7.1.2 lets a server put a trailer section
there and asks recipients to discard fields they do not understand;
instead the whole message failed and the resource never reached the
mirror.
The trailer section is now read the way headers are, as a block ending
on a blank line, and thrown away. Reading it as a block also bounds it:
trailers carry no length of their own, so an endless one would hold a
connection slot forever, and the line reader's 8KB buffer caps it. Only
the terminating chunk opens the section, so junk where a data chunk's
own CRLF belongs is still a framing error.
Closes#855
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0199wAkSVZNBNp51mpRkxMvv
Signed-off-by: Xavier Roche <roche@httrack.com>
* Renumber the trailer test to 149, 147 is taken by the WebDAV overflow test
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0199wAkSVZNBNp51mpRkxMvv
Signed-off-by: Xavier Roche <roche@httrack.com>
* Name the line-block bound, and pin the trailer edge cases
Review follow-ups: the 8190-byte cap that bounds a trailer section was an
unnamed literal inside http_xfread1, so raising it for large response
headers would have moved the trailer bound silently. It is now
HTS_LINE_BLOCK_SIZE, named where the reader is declared.
The trailer path also no longer runs the chunk-size parse it then
discards, and eof.html pins the deliberate leniency the change
introduces: past a complete, length-verified body, a trailer section cut
before its blank line still lands. A body cut before the terminating
chunk stays refused.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0199wAkSVZNBNp51mpRkxMvv
Signed-off-by: Xavier Roche <roche@httrack.com>
* Name http_xfread1's read modes instead of passing bare 0, -1 and -2
http_xfread1() selects its read mode from the sign of bufl: a positive
value reads that many bytes, 0 stops at a blank line, -1 at the first LF,
-2 only reserves the buffer. Nothing declared them, so every call site was
an unexplained literal.
Declare HTS_XFREAD_LINE_BLOCK, HTS_XFREAD_LINE and HTS_XFREAD_RESERVE in
htslib.h beside HTS_LINE_BLOCK_SIZE, and use them at each call site. The
selftest keeps its 8192, a byte count rather than a mode.
Behaviour-preserving: the preprocessed output of the changed translation
units is token-identical once the parens around the negative literals and
the __LINE__ digits shifted by the reworded comments are normalized away.
Closes#914
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Note that the reserve-only read mode has no caller
Naming it made it read as a supported mode; it is unreachable (#923).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0199wAkSVZNBNp51mpRkxMvv
Signed-off-by: Xavier Roche <roche@httrack.com>
* Keep the reserve-mode note on one line
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0199wAkSVZNBNp51mpRkxMvv
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Bound the ProxyTrack DAV item buffer against an amplified PROPFIND path
proxytrack_add_DAV_Item() reserved a fixed 1024 bytes and then sprintf'd into
it unbounded. The request path lands in the response twice, once as the href
and once as the displayname, and escapexml() turns each '&' into '&', so
an unauthenticated PROPFIND of roughly 900 ampersands writes about 9000 bytes
off the end of the heap block. No cache entry and no Depth: 1 are needed.
Replace the hand-sized reserve with StringSprintf(), which measures the
formatted output and grows the String to fit, and convert the sibling sprintf
sites in the same file so no unbounded write into a String is left to
re-audit. Sizing beats clipping here: the String already owns a growable
buffer, so nothing has to be dropped.
Closes#836
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Bound StringSprintf's pre-C99 retry, and trim the review findings
A genuine vsnprintf conversion error returns -1 just as pre-C99 msvcrt does
for a short buffer, so the doubling search had no way to tell them apart and
grew until realloc aborted. Unreachable from these format strings, which use
only %s and %d, but the helper lives in a shared header and will get more
callers. Cap the search and empty the String past it.
Also: the count assertion piped into wc under pipefail, so a zero count killed
the test through set -e before its diagnostic could print.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Lift the NO_WEBDAV conditional out of a macro argument list
A preprocessor directive inside a macro invocation's arguments is undefined:
it was fine while this was a plain sprintf() call, and MSVC rejected it as
soon as it became StringSprintf(). GCC accepts it, so only the Windows leg
caught it. Compute the DAV header fragment first and pass it as an argument.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Cover StringSprintf's exact-fill case and the WebDAV enumeration branch
StringSprintf_ writes the terminator at buffer[ret], so widening its
`ret < capacity` guard by one byte is a heap overflow that only fires when the
formatted output exactly fills the capacity. No crawl test lands on a
capacity boundary, so the mutant survived the suite. The new `strsprintf`
self-test sweeps lengths around 256, 512, 1024 and 2048 with the String's
capacity pinned to each, plus a growing and shrinking sweep on one reused
String, and checks the length, the bytes and the terminator every time.
Test 147 only ever sent Depth: 0, leaving the enumeration branch the same PR
rewrote with no coverage at all. Its fixture gains a child directory, and a
Depth: 1 listing pins the item URLs, including the trailing '/' that
StringPopRight takes back off a directory name.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Declare the new WebDAV test in the Windows skip set
It skips on Windows for the same reason as its two neighbours, MSYS
cannot reap a background listener (#595), and the ratchet fails a skip
it was not told about.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0199wAkSVZNBNp51mpRkxMvv
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* A chunked response carrying trailers is discarded as "Invalid chunk"
The chunk automaton expected the line after the terminating zero-length
chunk to be empty. RFC 9112 7.1.2 lets a server put a trailer section
there and asks recipients to discard fields they do not understand;
instead the whole message failed and the resource never reached the
mirror.
The trailer section is now read the way headers are, as a block ending
on a blank line, and thrown away. Reading it as a block also bounds it:
trailers carry no length of their own, so an endless one would hold a
connection slot forever, and the line reader's 8KB buffer caps it. Only
the terminating chunk opens the section, so junk where a data chunk's
own CRLF belongs is still a framing error.
Closes#855
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0199wAkSVZNBNp51mpRkxMvv
Signed-off-by: Xavier Roche <roche@httrack.com>
* Renumber the trailer test to 149, 147 is taken by the WebDAV overflow test
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0199wAkSVZNBNp51mpRkxMvv
Signed-off-by: Xavier Roche <roche@httrack.com>
* Name the line-block bound, and pin the trailer edge cases
Review follow-ups: the 8190-byte cap that bounds a trailer section was an
unnamed literal inside http_xfread1, so raising it for large response
headers would have moved the trailer bound silently. It is now
HTS_LINE_BLOCK_SIZE, named where the reader is declared.
The trailer path also no longer runs the chunk-size parse it then
discards, and eof.html pins the deliberate leniency the change
introduces: past a complete, length-verified body, a trailer section cut
before its blank line still lands. A body cut before the terminating
chunk stays refused.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0199wAkSVZNBNp51mpRkxMvv
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Spool a frozen backlog slot outside the mirror namespace
back_cleanup_background() named the spool file by appending ".tmp" to the save
name, so it landed beside the mirrored file. That is the shape #774 fixed for
the re-fetch backup: a site serving <path>.tmp has its mirrored copy truncated
by filecreate() and then unlinked when the slot is woken, and the run still
reports success. It is reachable on defaults, not only under a saturated
backlog: a -Z crawl of the bundled bigcrawl site with -c4 logs slots moving to
background.
Route both name shapes through back_spoolname(), which puts them in the
~hts-tmp directory no save name can spell, and drop that directory at the two
sites that unlink a spool. Left out of the #774 PR because these lines also
carried the overflow in #857.
Closes#859
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Keep the -p0 spool relative when no output directory is set
The new name inserted its own separator before ~hts-tmp, but path_html_utf8
already carries one and is empty when -O is absent, so the spool became
/~hts-tmp/tmpfile0.tmp: absolute, in the filesystem root. Master built
"%stmpfile%d.tmp" and stayed relative to the working directory.
create_back_tmpfile() has spelled it the same way since #842. Its empty
path branch looks unreachable from the three call sites, so this side is a
consistency fix with no test behind it, unlike the spool.
The self-test pinned the doubled slash it observed rather than the shape it
wanted; it now asserts the single-separator form and covers the empty
path_html_utf8 case that produced the root path.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The panel decoration was drawn as vector and flattened to a 4KB indexed GIF
some time around 2007, with the panel colour baked in as an opaque backdrop.
Refitting its four ellipse boundaries recovers the original geometry, so it
goes back to being what it was: two elliptical annuli, 553 bytes of SVG, with
a transparent background that now composites over the panel instead of having
to match it.
Rasterised at the same size, the only pixels that differ from the GIF are
single-pixel anti-aliasing fringes along the four boundaries. Nothing survives
a 3x3 erosion of that difference, so no edge has moved.
Dark mode still drops the image rather than inverting it, since the ring
lavender is a light-panel tone whichever way the file stores it.
The engine keeps its own embedded copy of this GIF for the backblue.gif it
writes into mirrors. That one is untouched: the filename and byte length are
a contract with pages already on disk.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* configure discards a user-supplied BASH_SHELL
AS_UNSET erased the variable before AC_PATH_PROGS could honour it, so
"./configure BASH_SHELL=/path" had no effect and there was no way to
point the build at a bash other than the first one on PATH. Nothing
presets BASH_SHELL, which was the whole problem with BASH in #895, so
declaring it precious is enough.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Run the nested configure against a symlink farm
An in-tree build leaves a config.status in srcdir, and autoconf then refuses
the out-of-tree run the test needs. Every CI build leg builds in-tree, so the
check failed there while passing on an out-of-tree tree.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Read the resolved bash from the configure trace, not the Makefile
The nested configure ran without the flags the outer one was given, so on
macOS it died at the openssl check that Homebrew paths satisfy. BASH_SHELL
is resolved long before that, so assert on the trace and let the run fail.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Assert the value that reaches $(BASH_SHELL), not the macro's decision
Reading the configure trace let a mutant through: resolve the override
correctly, clobber BASH_SHELL one line later, and both assertions passed
while every Makefile got the wrong shell. Prefer the generated Makefile
and keep the trace only as a fallback for a configure that dies early.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* The AppStream metainfo still advertises WebHTTrack 3.49.8
The metainfo installs to usr/share/metainfo, so GNOME Software and KDE
Discover read both the version and the "What's new" text out of it. Its
releases block held one entry, 3.49.8, and nothing had moved it since.
It now lists 3.49.8 through 3.49.15, newest first, each with a short
user-facing note taken from history.txt. Dates come from the git tags;
that also corrects 3.49.8's, which carried 3.49.7's date.
01_engine-version-macros.test gains two assertions so the next release
cannot miss this file, or configure.ac: AC_INIT and the top release entry
must both match HTTRACK_VERSIONID, and the entries must descend so the
top one really is the newest.
Closes#884
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Harden the metainfo version check and correct four release notes
The release-version extraction matched <releases ...> as well as <release>,
ignored XML comments and took only the last tag on a shared line, so a parked
or wrapper version= could pose as the newest entry and pass a stale metainfo.
Split tags one per line, drop comments, and anchor the match. Widen the awk
ordering key so a component of 1000 or more cannot borrow into the next.
In the notes: "3.49-2" and "site rules with wildcards" are unreadable in a
software centre, the Windows path bullet does not apply to the Unix WebHTTrack
GUI it ships with, and 3.49.15 listed no web-interface fix at all.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Find the data directory instead of trusting the configure-time one
webhttrack probed a fixed list of prefixes that nothing derived from
--datadir, and the engine baked $(datadir) into the binary with the
argv[0] fallback compiled out. Both fail on any tree that is not where
it was configured.
configure substitutes the real datadir into src/webhttrack, and
hts_resolve_datadir() prefers the compiled-in path but derives one from
argv[0] when it is gone, so a moved install reads its own templates
rather than silently falling back to the built-in defaults.
Closes#887Closes#894
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Resolve the data directory from the executable, not just argv[0]
MSVC broke: HTS_HTTRACKDIR is only defined on non-Windows, so the
argv[0] branch this replaced was live there, not dead. Windows now
passes an empty builtin and falls back to the executable's own
directory, which is what it did before -- except fconcat inserts no
separator, so the old path_bin lacked its trailing slash and never
resolved a template anyway.
Ask the OS for the executable path (/proc/self/exe, _NSGetExecutablePath,
GetModuleFileName) and keep argv[0] as the fallback, so a mirror run
through a PATH lookup resolves too.
The bundle drops the substituted datadir from its copy of webhttrack:
it is a build-machine path there, and the relative entries ahead of it
already find the payload.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Clip the candidate path instead of aborting on a long argv[0]
strlncatbuff() aborts rather than truncates, and appending the layout
suffix to an already-full buffer reaches that: a directory part within
17 bytes of the candidate buffer's size killed the process. Build the
candidate with snprintf and skip it when it does not fit.
The self-test now drives a directory part long enough to trigger it;
without the fix it aborts on "overflow while appending 'layout[i]'".
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Ignore the generated src/webhttrack
An in-tree build writes it next to webhttrack.in, where it was untracked
and one "git add -A" away from re-entering the tree with a build
machine's datadir frozen into it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* configure resolves bash to /bin/sh on macOS, and make deb runs a bash script with dash
AC_PATH_PROGS searched into BASH, which bash presets to its own invocation
path. configure re-execs through /bin/sh, and on macOS that shell is a bash, so
the macro honoured the pre-set value and reported "checking for bash...
/bin/sh": a bash in sh-mode that rejects process substitution. Search into
BASH_SHELL instead, a name no shell presets, with AS_UNSET in front so the
environment cannot preset it either.
That makes the obvious fix for the deb target safe. It ran tools/mkdeb.sh with
$(SHELL), which is /bin/sh, so "make deb" died on the first bashism on every
Debian and Ubuntu box. macos-app.sh stays on $(SHELL): it is POSIX sh on
purpose, so shellcheck lints it as sh.
tests/146_bash-shell.test asserts the configured shell exists, sets
BASH_VERSION, is not in POSIX sh-mode, and parses a process substitution.
Closes#895Closes#891
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Trim the comments this branch added, and drop a claim that is no longer true
The four comment blocks the branch added ran two to five lines where one or two
carry the fact. The 146 header also said the macOS shell rejects "the process
substitution the bundle script uses": tools/macos-app.sh has been POSIX sh with
no process substitution since #890, so the gate is there for tests/local-crawl.sh
and tests/webhttrack-smoke.sh, which is what the comment now says.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* The test server's bind reverse-resolves 127.0.0.1, which stalls on macOS
http.server's HTTPServer.server_bind() calls getfqdn() on the bind address just
to fill server_name, which nothing in local-server.py reads. On the macos-15
runner that lookup takes ~30s, so the PORT line lands well past every caller's
discovery budget (#870).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* PROBE: macos-15 + startup timing (not for merge)
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Move the macOS CI legs to macos-15, and route the last two port waits through the shared helper
84 and 100 kept their own PORT poll loops with 10s budgets, the copies #869 did
not reach; both use discover_server_port now, so there is one implementation
left.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Cut the server_bind comment to the why
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* armhf crash reports have no frames, and the empty backtrace fails the build
gcc emits no unwind tables on armhf, so backtrace() comes back empty and the
handler printed a frameless report, which the crash tests read as a failure and
which left 3.49.15-1 stuck at Build-Attempted there. Ask for
-fasynchronous-unwind-tables where the compiler takes it, and say why the report
has no frames when the unwinder still returns nothing.
Closes#892
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Let ASan tolerate the backtrace shim's link order
An LD_PRELOAD library loads ahead of the executable's own libasan, which ASan
refuses by default, so the sanitize leg never reached the crash it was meant to
inspect. Same waiver the other interposer tests carry.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Keep an empty backtrace a failure everywhere but 32-bit ARM
Relaxing test 80's skip to the shared message prefix let it swallow the new
"unwinding failed" wording too, so a build that traced nothing anywhere would
have gone green on every architecture. Skip only where nothing can be done
about it: the OS-less case, and 32-bit ARM if its toolchain still refuses to
unwind. Test 143 now pins the exact wording rather than the prefix it shares
with the OS-less notice, and its control run skips the symbolizer it has no
reason to spawn.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Declare the new backtrace test's Windows skip
The Windows job compares the skip set exactly, so a test that skips there for a
good reason still fails the gate until it is named.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Ship a macOS app bundle
macOS users install HTTrack through Homebrew and get a working WebHTTrack they
are never told about: the formula installs webhttrack and htsserver, and the
only thing missing is something to double-click. tools/macos-app.sh assembles
HTTrack.app from an installed prefix, with the payload under Contents/Resources
so webhttrack keeps resolving htsserver and its data from its own location, and
a two-line stub in Contents/MacOS for Launch Services.
Nothing about the engine changes. The bundle is possible because webhttrack was
already relocatable and because the data symlink stopped being absolute (#885);
--disable-shared keeps libhttrack inside the binaries so nothing points back at
the staging prefix. That costs no crash diagnostics here, since backtraces are
gated on __linux (src/htsbacktrace.c:50), though it does trip #889 on Linux.
The script verifies what it builds rather than trusting it: no absolute symlink,
the served UI present, no Mach-O still linking the staging prefix, and the
Info.plist version matching the installed binary. configure generates that
plist, so it cannot drift into a fifth hand-maintained version spot of the kind
#884 describes.
CI assembles the bundle, runs the webhttrack smoke through the stub, then moves
the bundle and deletes the prefix it came from and runs it again, which is the
one thing a .app has to survive that a prefix install does not. An ad-hoc
codesign proves it is well formed enough to sign; Gatekeeper needs a Developer
ID and stays out of scope with the DMG.
No custom icon: the largest artwork in the tree is 48x48 and macOS wants 1024,
so that needs a real source asset.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* ci: lint the new bundle script, and mark it executable
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Drive the bundle through a make target, and run it with bash
Adds a macos-app target so assembling the bundle goes through the build system
the way make deb does, rather than CI reaching for the script directly.
It runs the script with $(BASH), not $(SHELL). automake's SHELL is /bin/sh, and
the script uses process substitution, so under dash it died partway: the payload
was already copied by then and only the verification was skipped, leaving a
bundle that looked built and had been checked by nothing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Make the bundle checks catch what they were missing
The smoke put the bundle's own bin on $PATH, so a launcher that ignored its own
location and just ran "webhttrack" passed as readily as the real one, which is
the most likely way a stub breaks. Dropping $prefix/bin from $PATH kills that:
the browser stub is found through webhttrack's SRCHPATH, not $PATH, so nothing
else needed it.
The bundle also shipped libtool .la files, static archives and include/, none of
them loadable from a static build and the .la files carrying the staging prefix
in libdir=. They are pruned, and a text sweep now fails on any remaining file
that embeds that prefix, which otool cannot see because it reads load commands
only. The prefix is resolved to an absolute path first, or a relative --prefix
made that grep match nothing.
Also: the symlink scan asserts it scanned something, and CFBundleVersion is
compared as well as CFBundleShortVersionString.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Assemble the bundle with POSIX sh, not $(BASH)
$(BASH) is not a reliable bash on macOS. /bin/sh there is bash in sh-mode, which
presets $BASH to the path it was invoked as, and AC_PATH_PROGS honours a
pre-set value rather than searching, so configure reports "checking for bash...
/bin/sh". That shell rejects process substitution, and the bundle job died on
it.
Rather than hunt for a real bash, the script no longer needs one: the three
process substitutions become temp-file loops, pipefail goes (not POSIX), and
the target is back on the ordinary $(SHELL) like deb:. shellcheck now reads it
as sh, so a bashism creeping back fails lint instead of macOS CI.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Make the installed html symlink relative
The install-data-hook linked share/httrack/html to an absolute $(htmldir), so
an installed tree only worked at the prefix it was configured for: the link
dangled under DESTDIR staging and in any relocated copy, and webhttrack then
failed its test -d "${DISTPATH}/html" check and exited. It now emits a
relative link when htmldir sits under datadir.
The hook also sat in html/Makefile.am while writing into $(datadir)/httrack,
the directory lang/Makefile.am declares and populates, and it hardcoded
$(prefix)/share instead of following datadir. Moving it to lang/ with
$(langrootdir) drops that cross-subdirectory install-order dependency. A
stale symlink was previously left in place, so an upgrade kept an absolute
one; it is now replaced, a real directory is left alone with a diagnostic,
and a new uninstall-hook removes what install created. A moved datadir still
defeats webhttrack's own data search for an unrelated reason (#887).
The smoke test asserts no installed symlink is absolute and that the link,
resolved inside a copy of the tree at another path, stays inside that copy
and lands on the served UI. It now runs on Linux as well as macOS, since
Linux is where the affected packaging is consumed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* ci: match the Linux build job's package list
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: put the stub browser where webhttrack looks first
webhttrack searches its own SRCHPATH (starting at $BINWD, then /usr/local/bin,
/usr/share/bin, /usr/bin) and only appends $PATH after it, so shadowing
x-www-browser through PATH only worked on hosts that have no real one. The
GitHub Linux runners ship /usr/bin/x-www-browser as Edge, which won the search
and then aborted on its SUID sandbox, failing the new Linux job.
Write the stub to $prefix/bin instead, which is SRCHPATH[0] on every platform,
and remove it in the teardown.
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Version bump to 3.49.15 in the three spots that carry it (configure.ac, htsglobal.h, version.rc), plus the release notes and the Debian changelog entry.
VERSION_INFO goes 3:6:0 to 3:7:0, soname still .so.3. The installed htsblk gained a field at its tail and lien_back embeds it by value, so lien_back.is_update and everything after it shifts 8 bytes. An offsetof probe against both tags confirms httrackp's own fields did not move. Same call as 3.49.14: HTTrackQt is the only consumer of these headers, so no libhttrack4 rename.
109 commits since 3.49.14, none of them under debian/, so the changelog entry is a plain new-upstream one with a note that httrack-doc grows about 1 MB from the new GUI screenshots.
The lede led with WARC output and the change report, which are niche features, because those happened to be the last things merged. What matters to someone arriving now is that HTTPS works, that proxies work, that a current page with responsive image sets and lazy loading captures properly, and that an update does not re-fetch the whole site.
The rings the 2007 pages carried behind their content panel come back, and now on every page rather than the handful that had them. The image has the light panel colour baked in, so dark mode drops it rather than showing a pale rectangle.
While in there: the wordmark is black on transparent and all but vanished against the dark field, which every page inherited along with the shared chrome. It is inverted in dark mode now.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
addurl.html was 116 words and five screenshots of a 2003 dialog, reached only from a single sentence in the guide. Its substance moves into that sentence: what the Add a URL button does that the address box cannot, including the browser-capture proxy, which lang.def confirms is still in the GUI. The stale screenshots go with it, and a fresh capture of that dialog belongs in the next screenshot walk rather than in a page of its own.
options.html had been an orphan since #661 gave it a legacy-list banner and repointed everything that used to link it. It becomes a stub pointing at the generated option reference and the command-line guide.
plug_330.html documents the callback interface of releases 3.30 to 3.40. Its title said so; now a banner does too, the way Fred Cohen's guide is labelled on the index, because someone arriving from a search engine reads the page and not the title.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
Nine of the twenty pages linked nowhere at all and seven had a single inbound link, so a reader landing on cache.html from a search engine got no navigation, no breadcrumb and no way back. Each page also carried its own copy of the same inline stylesheet inside a six-table shell with a 400px floor, which is why sixteen of them scrolled sideways on a phone.
They now share the guide's chrome: masthead, sidebar, footer, and where the headings allow it a generated "On this page" list. file:// has no server includes, so the navigation is real markup in every page, written by tools/doc-chrome.py and verified by CI; tools/doc-links.py resolves every relative link and fragment. Both refuse to pass vacuously, and both were checked against planted defects.
The prose is untouched. For twelve of the fifteen pages the content text is identical word for word; the three exceptions are cmdguide.html losing the contents list the sidebar now carries, contact.html re-encoded to UTF-8, and one font tag whose removal joined two words the browser already ran together. Each page also gains a real title and its own description, replacing the shared blurb whose keyword list still advertised Windows 95 and AIX 4.0.
Image zoom and sidebar highlighting move to doc.js so every page has them, leaving guide.js the platform switcher and the option filter. The guide renders identically at 1100px in light and dark, and thirteen behaviours pass unchanged before and after; at 390px one Proxy paragraph that used to overflow its container by 37px now wraps. Nothing scrolls sideways at 390px or 1100px, and no page throws a script error.
httrack.man.html keeps its own bare styling for now, since reskinning it belongs in the generator that writes it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
The documentation index is the page WinHTTrack opens for generic Help and the one httrack.com links as "Documentation", and it presented twenty destinations as a flat bullet list with nothing to say which one a newcomer wants. It is now a task-grouped hub in the interface guide's idiom, with overview.html folded in as the lede and refreshed for 3.50, and Fred Cohen's guide under an Archive heading labelled as written for 3.10, keeping its URL.
The chrome the hub needs is split out of guide.css into doc.css; guide.css keeps the platform switcher, tab strip and option entries. guide.html renders pixel-identically before and after the split at 1100px and 390px, in light and dark.
overview.html, start.html and cmddoc.html become meta-refresh stubs, like the retired step pages. start.html was a window.open popup launcher and cmddoc.html has 404ed since #661 deleted it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
Test 142 guards against a vacuous run by requiring every asset it types to exist. img/android_spider.png stopped existing when the Android help page was folded into the unified guide and its screenshots were renamed guide-droid-*.png: #874 and #876 were in flight at the same time, so neither run saw the other half, and master has failed this test on every platform since it landed.
The three references now point at guide-droid-opt-spider.png, a PNG in the same directory, so the check still tests what it was written to test.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* htsserver labels every PNG as image/gif, and serves JPEG as a download
The response header was picked per broad file class, so is_image() matched
both .gif and .png but always emitted Content-type: image/gif, and anything
outside the five hardcoded classes (.jpg, .svg, .ico, .xml) went out as
application/octet-stream, which browsers download instead of render.
Replace the per-class headers with an extension to type table, matched on the
real extension rather than on a substring of the path.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Cover the untested paths the content-type table can regress on
The first test named 9 of the 12 extensions and no superstring case, so a
lookup matching on a prefix, or one that dropped the htm or jpeg row, stayed
green. Add those fixtures, a raw-socket check that the header line is still
CRLF-framed, and the no-cache pair that only the HTML branch carries.
Also serve .xml as application/xml (RFC 7303 deprecates text/xml) and add
webp and avif, the two image types a browsed mirror most often needs.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* One guide for WinHTTrack, WebHTTrack and Android, instead of 18 thin pages
The GUI documentation was 18 files whose bodies were mostly a copy-pasted style
block, illustrated with 1998-era screenshots of WinHTTrack, with the Android app
documented separately and WebHTTrack not shown at all. html/guide.html replaces
them with a single page: one walkthrough and one option reference, and a
switcher that swaps the screenshots and the platform-specific notes between the
three versions. Each option now also names its command-line equivalent.
The screenshots are the three capture sets from httrack-works, prepared by
tools/doc-images.py. Deep links carry the platform, so guide.html#win/opt-limits
opens the Limits section already showing the Windows dialog; the WebHTTrack
option pages use that form. The old filenames stay as redirect stubs for the
forum links that still point at them.
Scripting is an enhancement throughout: without it the page shows every
platform's text and screenshots rather than hiding any of it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Document the options the old pages never had
A coverage probe against the WebHTTrack form controls found the reference
inherited the old doc's blind spots: WARC, the CDXJ index and WACZ, --changes,
sitemap seeding, cookies-file, the urlhack opt-outs, strip-query, the proxy type
selector and the random inter-file pause were all exposed in the UI and absent
here. Labels and descriptions taken from lang.def, CLI flags from htsalias.c.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Add tools/screenshot-walk.py, which starts htsserver, drives its pages in
headless Chromium and captures one PNG per documentation screen, plus a
manual screenshots workflow that runs it on a runner. The Android and
Windows front ends already have their equivalent.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Refresh the Android help page from current app screenshots
The ten screenshots dated from an older build and covered only two of the
eleven option tabs. Recaptured all of them from app versionCode 85 (engine
3.49.14) and added Build, Spider and Log/Index/Cache, whose contents the
one-line summaries could not convey.
The new shots also correct three claims: the proxy tab now picks a protocol
(HTTP, SOCKS5, HTTP CONNECT tunnel), Log/Index/Cache gained WARC output and
the change report, and Base path opens a chooser rather than being fixed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Document every option tab on the Android page
Android users cannot easily flip between help pages, so the page now carries
its own option reference: all eleven tabs, each with its screenshot and every
setting on it, replacing the pointer at the desktop reference.
Six settings the desktop pages never covered are documented here for the
first time: keep-alive, the inter-file pause, the three URL-hack opt-outs on
Links, and Strip query keys on Experts Only.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
GitHub marks macos-14 deprecated, so both macOS jobs are on an image headed
for brownouts and removal. macos-15 is the same arm64 host and every path in
those jobs is brew --prefix-derived, so nothing else moves.
Also point Dependabot at git submodules: src/coucal is a hard build dependency
that only advances when someone remembers to look.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
local-crawl.sh waits 30s for local-server.py's "PORT <n>" line and matches it
anywhere, because a cold Python start under a parallel make check lags past a
second and a warning merged via 2>&1 can precede the line. Seventeen tests that
launch the server directly carried an older copy of that loop: 5s, and only the
first line of the log. macos-15 exposed it, failing 15 of them where macos-14
passed on margin, but nothing about the bug is macOS-specific.
Move the proven idiom into testlib.sh next to stop_server and call it from all
seventeen, which also gets them the log dump on failure that local-crawl.sh has
and the inline copies did not -- the reason the CI failure said only "could not
discover server port" with no way to tell a slow start from a dead server.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
OpenSSL 4.0 removes what 3.x deprecated. The engine is already clean against
that set, but nothing holds it there: a deprecated call would build fine on
every current leg and only surface when a distro moves to 4.0.
Build with -DOPENSSL_NO_DEPRECATED, which tracks whatever OpenSSL the runner
has rather than pinning a level the headers may reject. Compile-only; the
header contract is what is under test and the matrix already runs the suite.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Add a hidden -#c option that crashes on purpose, for crash-handler testing
The crash handlers had no way to be exercised: httrack's own fatal-signal
handler, WinHTTrack's, and the Android app's mirror-recovery path could only
be tested by waiting for a real bug. -#c[=KIND] faults deliberately, in
segv, abort, trap or stack flavours, announcing on stderr that the crash was
asked for so a log never reads as a genuine failure.
Undocumented and unlisted, but always compiled in: the point is to test the
shipped artifact (the APK, the installer, the distro package), which a
build-time flag would exclude.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Assert a backtrace portably: macOS has no backtrace() at all
The frame check matched glibc's "[0xaddr]" form, which macOS never prints:
USES_BACKTRACE is gated on __linux, so every other OS gets the "No stack
trace available" notice instead. Accept either, and also require the
handler's closing line so the assertion still proves it ran to the end.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* The webhttrack GUI cannot reach --warc-cdx, --wacz or --warc-max-size
PR #672 exposed --warc and --warc-file on the "Log, index, cache" tab; the
three WARC flags added after it never got a control, so a CDXJ index or a
WACZ package is unreachable from the GUI. Add the two checkboxes and the
size field, and translate the six new lang.def strings into all 30
languages instead of letting them fall back to English.
An unknown ${LANG_*} renders as the empty string, so a key missing from
lang.def or from one lang/*.txt blanks a label instead of failing: test 140
now asserts every key the templates use exists everywhere.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Move the lang completeness check into 62_lang-integrity, its owner
62_lang-integrity.test already owns lang.def/lang-file consistency; it
checked that no msgid drifted, but not that every msgid is translated
everywhere, which is the gap that let the WARC options ship untranslated.
Add the completeness pass there and drop the separate 140 test, which
reimplemented the same parser in Python and was scoped narrowly enough to
be blind to the 24 msgids already missing from 26 files. Those are waived
by name against #863, and a waiver that stops being needed now fails.
Also check every ${LANG_*} the templates interpolate resolves, in any
wrapper form: the ${html:html:...} and ${fexist:...:...} spellings were
outside the dropped test's regex.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* url_savename never clamps the final path segment, only the whole path
Inside url_savename's over-ceiling shortening, every directory segment is cut
to MAX_SEG_LEN but the name itself is copied under the whole-path budget alone.
A single component could therefore run to ~226 characters on Windows while its
parent directories were held to 48, and once the directories had eaten the
budget the cut landed mid-name and took the extension with it, leaving a page
saved with no .html at all.
Clamp the name like any other segment, and reserve a plain extension across the
cut the way #623 already reserves the ".delayed" marker.
75_engine-longpath-posix's Windows branch asserted the clamp all along; it had
never run until #847 broadened the CI glob, so this also drops the expected_skip
that filed the failure as #852.
Closes#852
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Fix 01_engine-savename's Windows expectation for the clamped name
The Windows arm pinned the pre-fix output: 210 characters of name and no
extension at all. It now takes the 48-char segment clamp with its ".html"
intact.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Percent-escape a query-string character reference the page charset cannot represent
On a page with no declared charset (iso-8859-1 fallback), a reference like
€ was left as literal source text in the query string: its '&' and '#'
then act as delimiters, so the origin parsed more parameters than the document
expressed. The URL Standard writes such a code point %26%23<decimal>%3B.
Adds hts_unescapeEntitiesWithCharsetSpecial() with an opt-in
UNESCAPE_ENTITIES_URL_QUERY flag, set only by the query-string call site; the
plain decoder keeps its text contract and its in-place guarantee. The query is
now decoded out-of-place, since the escape grows the string.
Closes#854
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Detect an unrepresentable code point with the encoder, not a round trip
The round-trip comparison escaped code points the charset does represent:
hts_convertStringToUTF8() short-circuits on an ASCII string and hands the bytes
back undecoded, so every reference on ISO-2022-JP, UTF-7 or HZ compared unequal,
and glibc's CJK tables are not injective (cp932 U+301C came back as U+FF5E).
On POSIX iconv already fails outright on a code point it cannot encode, so the
NULL return is the answer. Windows substitutes instead of failing, and reports
it through WideCharToMultiByte's lpUsedDefaultChar, which we now pass;
hts_convertStringFromUTF8Strict() turns that into a NULL. Test 139 grows a
declared shift_jis and a declared iso-2022-jp page, which the two charset cases
it had could not have caught.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Assert the ISO-2022-JP query per field, and let the two cp932 tables differ
Windows maps 0x8160 to U+FF5E rather than U+301C and reports the substitution,
so 〜 is genuinely unencodable there and the escape is right: accept either
encoding for the shift_jis page. ISO-2022-JP's trailing shift back to ASCII also
differs between iconv and Windows, so that page is asserted field by field, which
is what the round-trip check would have broken anyway.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Let Windows substitute the euro sign on a stateful codepage
WideCharToMultiByte refuses lpUsedDefaultChar on ISO-2022-JP, so the
substitution cannot be seen and € goes out as '?' there, as it did before this
branch. It splits no query, and the page is in the test for the references the
charset does represent.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Broaden Windows CI's engine-test glob and skip-set compare (#844)
The Windows engine-test loop named individual tests (01_engine-*.test plus a
hand-maintained tail) because its glob missed NNN_engine-*.test, so a new
test landed with zero Windows coverage until someone remembered to add it by
hand; two branches doing that at once collided on the same line. Broaden the
glob to *_engine-*.test, matching the *_local-*.test pattern already used
next to it, so a stray test file gets picked up automatically instead of
silently skipped.
That newly covers 66_engine-port80-strip, 67_engine-delayed-truncate,
75_engine-longpath-posix, 80_engine-crash-symbolize and
104_engine-warc-longurl, none of which had ever run on a Windows runner.
80_engine-crash-symbolize exits 77 there by design (no backtrace()), which
would have tripped the old exact-string expected_skips gate and turned the
leg red. Switch that gate to a sorted-set compare (one name per line, so an
append no longer shares a line either) and add the new skip; a mismatch now
prints a diff naming exactly which test's skip state changed instead of a
raw string dump.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Generalize the last two hand-pinned test names in the Windows CI loop
13_crawl_proxy_https.test and 60_crawl-log-salvage.test matched no
glob, so a renumber or a same-named future test would silently drop
Windows coverage again. Tightly globbed instead of using *_crawl*,
which would also sweep in the live-network crawl-*/crawl_https tests
this job never runs.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* kill_tree's Windows path falls through into a POSIX process-group kill
test-timeout.sh and run_with_timeout skip set -m on Windows (nothing there can
signal a group), so no target of kill_tree ever heads its own process group.
kill_tree's taskkill branch already reaps the Windows tree, but had no return,
so it always fell through to kill -9 -"$pid": on Windows that treats $pid as a
process-group id it never is, and can hit whatever real group its number
collides with under MSYS's own pid churn -- including the harness's own,
which is how PR #847's CI run died silently right after 102_local-ftp-refetch
finished, mid-cleanup, with no FAIL line and a truncated artifact.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <xroche@gmail.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* A chunked response cut at a chunk boundary is stored and cached as complete
back_finalize()'s incomplete-transfer guard is gated on r.totalsize >= 0, so it
only fires for a response that declared a Content-Length. A chunked response
keeps totalsize == -1 until the chunk headers add up, and the sum always equals
what arrived, so a stream that ended before its terminating zero-length chunk
took the success branch: the partial body was mirrored, cached, and on an
--update replaced the previous good copy.
The terminating chunk is part of the framing, so its absence makes the body
truncated by definition. Track it through chunk_blocksize, which only reaches
-1 once that header is parsed, and refuse the transfer in the two places the
short-Content-Length path already covers: back_finalize skips the cache, and
back_wait fails the transfer so the engine retries and leaves the mirrored copy
alone.
An HTTP/1.0 identity body delimited only by connection close has no framing to
check, and is deliberately left alone.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Review fixes: cache assertions, -%B coverage, disk-write limit
Adds the cache half of the spec the test never asserted (--cache-found /
--cache-not-found in local-crawl.sh, reading hts-cache/new.zip), a -%B leg
covering opt->tolerant on a chunked body, and a direct-to-disk .bin resource.
Gates the back_wait block on r.statuscode > 0 like its back_finalize
counterpart, and drops the chunk_size == 0 hunk: a negative chunk-size header
is already refused by the short-body guard and the Invalid chunk path.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Restore the chunk_size == 0 sentinel guard, and pin it
A chunk-size line of 80000000 lands in an int as INT_MIN via sscanf("%x"):
a bare else claimed it as the terminating chunk while totalsize went negative,
so neither the framing check nor the short-body check fired and an --update
deleted the good copy and cached a 0-byte 200. Invalid chunk cannot cover this,
being set after back_finalize has already stored the response.
The update pass of chunktrunc/hostile.html now serves that header, and
chunktrunc/reset.bin aborts its body with an RST so a read error keeps its own
diagnostic instead of the framing check relabelling it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Bound the parsed chunk size before it reaches the in-RAM realloc
sscanf("%x") into an int lands a chunk-size line of 80000000 on INT_MIN, which
drives totalsize negative past the INT32_MAX guard and sign-extends into a 16 EB
realloc (#860, pre-existing on master; the new #840 test is what reaches it).
Parse the field with strtoull and drop anything an int cannot hold, so the
existing arithmetic and the in-RAM cap keep working on a value known to be in
[0, INT32_MAX].
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <xroche@gmail.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* A URL can still be saved onto the engine's own temporary files
The re-fetch backup and the content-coding temporaries built their names
inside the mirror, so a site serving the matching path had its mirrored
copy used as scratch and then deleted. #818 escaped hts-cache and hts-tmp
as reserved path components, but three shapes still got through (#842): a
trailing space that cleanEndingSpaceOrDot() stripped back off after the
check, and the first path component, which the table could not see because
it anchors on a leading '/' that url_savename has already removed. The
latter covers a single-label hostname too, and left the DOS device names
equally exposed in that position.
The temporaries now live in ~hts-tmp/. url_savename maps '~' to '_' in
every name it builds, so nothing a site can serve resolves inside it
whatever the escape does; the escape is still fixed, as defence for
hts-cache/ itself. The frozen-slot spool (<save>.tmp) shared the
collision, plus a heap buffer sized from url_sav that the path_html branch
overran, and moves to the same helper. The .delayed placeholder does not:
it goes through url_savename's collision detection, so a competing link
gets a suffix rather than the file.
Closes#774Closes#842
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Scope the reserved-name escape and fix the review blockers
Restrict the new first-component match so a hostname label is not rewritten:
only a trailing run cleanEndingSpaceOrDot() strips may end the name there,
so aux.example.com keeps its name while a single-label hts-tmp host and a
path-first hts-cache/ are still escaped. Renaming a host directory would
have moved an existing mirror out from under --update --purge-old.
Drop the frozen-slot spool rewrite: it carried a pre-existing heap overflow
that belongs in its own change. Clear back->tmpfile when structcheck() fails
in the named branch, as the unnamed one does. Test 132 now plants a leftover
at the exact temporary path, so both a reverted HTS_TMPDIR and a no-op
back_refetch_backup fail it; 131 pins the hostname labels.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* The frozen-slot spool name is written into a buffer sized from a different string
back_cleanup_background() sized the temporary's name buffer from url_sav and
then, under -p0, wrote a path_html-derived name into it. Build the name in a
bounded buffer and duplicate it, so neither form can overrun.
Closes#857
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Assert the spool actually ran, and correct the overflow's blast radius
The -p0 name escapes on a plain build too: _FORTIFY_SOURCE=3 catches the
same write in __sprintf_chk, so the abort is not sanitizer-only. Test 138
now runs under -Z and requires the engine log to show a slot spooled and no
serialize error, so a build with neither fortify nor ASan still proves the
branch was taken. tmpname gains room for the ".tmp" a full-length save name
appends, so no input is refused that the old exact-sized buffer accepted.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* A 304 during --update leaks the whole previous htsblk
back_wait() handles a 304 by replacing the response struct with the cache
entry, carrying only the socket and keep-alive members across via
back_connxfr(). The struct assignment drops every owned pointer the live
response still held without freeing any of them: the 8 KB header buffer on
every update, plus the two WARC header stashes when --warc-file is on. An
update over a 10k-page site drops roughly 80 MB in one run.
back_clear_entry() already knew how to tear those down, so the frees move into
a helper that both it and the 304 path call.
The new test runs the two-pass mini304 crawl with LeakSanitizer on, which the
sanitized CI job otherwise disables. The fresh first pass is the control: it
has no cache entry to read back and is clean either way. The update pass
reports 16 KB in 2 objects on master, one per unchanged URL, and nothing with
the fix.
Closes#782
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Trim the test header
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Review fixes: reuse deleteaddr(), and cover the WARC limb
back_free_response() was reimplementing deleteaddr(), which already frees adr
and headers and NULLs both; call it instead so the two cannot drift.
Test 114 never passed --warc-file, so the warc_free_request() limb ran with
both pointers NULL on every path it exercised and deleting it kept the test
green. A third pass turns the archive on, and it now fails with the 835 and 238
byte stashes when that call goes away.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Register the new leak test as an expected Windows skip
The Win32/x64 job pins the exact set of tests allowed to skip, so an
all-skipped suite cannot report green. 114_local-update-304-leak needs a
LeakSanitizer build and MSVC has no equivalent, so it skips there and tripped
the gate with fail=0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* A 304 revisit archives none of the exchange it stands for
The cache entry overwrites the whole htsblk in back_wait(), taking the
stashed request and response headers with it, so the revisit record gets
a synthesized "HTTP/1.1 200 OK" block and no request record. Both blocks
now move across the swap, but only for a real 304: the engine also forces
HTTP_NOT_MODIFIED by itself, and those have no 304 to carry. Since a 304
declares no Content-Type, warc_write_transaction() takes the resource type
from the caller so revisit CDXJ lines keep their mime.
cache_read_including_broken() returns the entry's response struct after
back_clear_entry() has freed it; adr, headers and location come back NULL
now, as the normal cache_readex() path already returns them.
Closes#826
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Do not expand an empty bodyargs array under macOS bash 3.2
Test 122 asks local-crawl.sh only for the revisit checks, so it is the
first caller to leave bodyargs empty. Bash before 4.4 calls that unbound
under set -u, which passes on Linux and fails the macOS leg.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Restore the opt argument to hts_rename_over call sites
Lost when the master merge took this file wholesale; #816 gave the function an
httrackp parameter and the three call sites here reverted to the old form.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* A 304 revisit names no peer IP, and an engine-forced not-modified is archived as one the server never sent
back_connxfr() moves the socket and keep-alive members across the cache-entry
swap in the 304 path of back_wait(), but not the peer address, so every
revisit record came out with no WARC-IP-Address (#838). Carry it across the
same swap.
Separately, several engine hacks (the same-size heuristic, the #176
error-ignore-on-update path, ...) force HTTP_NOT_MODIFIED themselves when the
server actually said something else. That statuscode flows into the same 304
branch, and warc_write_backtransaction() turned it into a revisit record whose
WARC-Profile claims a server-not-modified 304 that was never exchanged (#839).
The fix reuses the existing server_sent_304 signal (already used to gate the
request/response header carry-over): its persisted proxy is
warc_resphdr != NULL, present only when the swap actually moved real 304
headers across. When it's absent, no server exchange happened this run to
archive, and there's no freshly re-fetched body either to dedup against a
same-run predecessor, so the fix skips the record entirely rather than
writing an identical-payload-digest revisit with nothing behind it to refer
to.
tests/134 and tests/135 extend the mini304/errmask fixtures into a two-pass
--warc-file crawl and assert against tests/warc-validate.py, which gained
--expect-ip and --no-record-for for these checks. Both reproduce their bug
when run against the pre-fix code.
Closes#838Closes#839
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Fix review findings: an engine-forced not-modified must still count as unbacked
Writing no WARC record at all for an engine-forced not-modified (the #839 fix
in the previous commit) dropped w->unbacked_revisits accounting along with it,
so an all-forced pass wrongly looked fully backed: warc_commit's swap guard let
it replace the previous archive with an (almost) empty one, reopening #777's
exact shape, and warc_cdx_flush then left a stale .cdx pointing into the
truncated file. Fix: emit a self-referencing identical-payload-digest revisit
instead of nothing, still counted as unbacked, so the guard keeps protecting
the previous archive exactly as it does for a real 304.
Classifying a notmodified htsblk by warc_resphdr's presence was too broad: it
also matched back_add()'s cache-priority branch, which never sent a request at
all and is unrelated to the back_wait() hacks #839 targets. Replaced it with an
explicit hts_boolean, warc_forced_notmodified, set only inside back_wait()'s
304-swap block from the already-computed server_sent_304, leaving every other
notmodified path (including back_add()'s) at its default and thus at its prior
behavior.
tests/134 now asserts the exact peer IP instead of mere presence, since a
zeroed-but-AF_INET address also satisfies a non-empty check. tests/135 gained
a second scenario: an all-forced-not-modified rerun against the same archive
name, which is the shape that lets the swap guard's bug actually destroy the
previous archive; the mixed-fixture scenario alone couldn't see it, since a
real 304 alongside it already keeps the guard's counter above zero.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Cover the cache-priority path and gate the digest revisit on OpenSSL
warc_forced_notmodified had exactly one assignment site, inside back_wait()'s
304-swap block. back_add()'s cache-priority branch (-C1, a lock-file resume,
or opt->state.stop) sets notmodified too but never sent a request, so it fell
through to the zero default and got classified as a genuine 304, emitting a
server-not-modified revisit around a synthesized "HTTP/1.1 200 OK" block. Set
the flag there as well: no request there ever means no server 304 to claim.
The identical-payload-digest revisit for an engine-forced not-modified also
assumed a digest was available, but payload_digest_b32() returns nothing
without OpenSSL. Gate the profile on have_pdig; write no record when it's
absent, since there's neither a digest to point a revisit at nor a real
exchange to record as a response, keeping w->unbacked_revisits incremented
either way so the archive-protection guard still holds.
tests/136 drives the cache-priority path with -C1 and checks the same
WARC-Profile assertion as tests/135. Both new branches (135 and 136) split on
HTTPS_SUPPORT so the no-OpenSSL CI leg exercises the no-digest path instead of
skipping it; 135's archive-kept scenario drops --wacz for the same reason, a
package needs OpenSSL for its SHA-256 digests and would otherwise skip the one
thing it exists to catch.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* An empty body looks digest-less too, and the profile tests never checked the real 304
has_payload requires body_len > 0, so a genuinely empty body (a real,
zero-byte file) fell into the same "no digest" branch as a build without
OpenSSL and got silently dropped from the archive instead of a proper
identical-payload-digest revisit. A zero-length body has a well-defined
digest (sha1 of nothing); compute it directly for that case instead of
reusing has_payload's gate.
tests/135 asserted mini304's pages were revisits but never which profile, so
forcing warc_forced_notmodified true unconditionally (a regression that would
mislabel or suppress a real 304) still passed. Added the server-not-modified
assertion for both mini304 pages, and a matching zero-length fixture
(errmask/empty.dat) plus its identical-payload-digest assertion. tests/136
gained a body-hex check on the fresh archive: its update-pass checks only
assert absence, so a warc_write_transaction stub that writes nothing at all
went undetected on the no-openssl leg.
The unbacked-revisits counter still increments even when nothing gets
written (load-bearing: it's what keeps the swap guard from letting an
all-unchanged pass overwrite a good archive with an empty one), so the
"this pass revisited N URL(s)" log line no longer claims those URLs got a
named body; it now says the archive doesn't hold their current content,
true whether a record was written or not.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
PR #794 moved four gmtime() call sites onto a reentrant hts_gmtime()
helper, but a handful of localtime()/gmtime() sites survived: the
progress/log timestamps in htsback.c and htscore.c, ProxyTrack's
timezone lookup in proxy/store.c (which also dereferenced localtime()'s
result with no NULL check), and the RFC822 date helpers in htslib.c.
glibc's gmtime/localtime buffer is process-wide, not per-thread, so two
of the engine's worker threads racing on it can each read back the
other's broken-down time.
Adds hts_localtime() next to hts_gmtime() and converts every remaining
call site.
Separately, time_gmt_rfc822() and get_filetime_rfc822() fell back to
localtime() whenever gmtime() failed, then handed the result to
time_rfc822(), which always appends "GMT" to the formatted string.
Local time labelled GMT is wrong regardless of which thread wins the
race, so both now fail (empty string / return 0) instead of guessing.
gmtime() essentially never fails on a valid time_t, so this path is not
reachable in practice; the new test guards it at the source level
instead of trying to trigger the failure at runtime.
tests/133_engine-reentrant-time.test drives a new -#test=localtime
self-test: a concurrent hts_localtime() stress test against a fixed
reference table (mirroring #794's hts_gmtime test, TZ pinned to a
non-UTC offset so a local/UTC mix-up cannot hide behind a UTC CI
runner), a get_filetime_rfc822() GMT round-trip check, and a source
guard confirming time_gmt_rfc822()/get_filetime_rfc822() never
reference localtime() again. Verified red on two independent mutants:
reverting hts_localtime() to a raw localtime()-and-copy corrupted
17549/400000 concurrent conversions, and reintroducing the local-time
fallback tripped the source guard.
Closes#806
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* A PROPFIND on an exact cache entry crashes proxytrack
PT_Enumerate() reports a folder's default document as a zero-length name,
and the WebDAV listing loop read thisUrl[thisUrlLen - 1] on it, four
gigabytes past the string. One unauthenticated request took the whole
listener down.
Closes#828
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Pin the new skip on Windows and tighten the listing assertions
The Windows job runs *_local-*.test and compares the skip list against an
exact string, so an unpinned skip fails the leg with fail=0, which reads
like a flake. Assert the href and the response count too: displayname
alone comes from the href's trailing component, so it cannot see a wrong
path above it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* The .bak re-fetch backup destroys a mirrored file of the same name
The re-fetch backup, the .z content-coding spool and its .u decode
target were all named by appending an extension to the local save name,
which is a name the mirror can hold. A site serving <path>.bak had its
copy renamed over as the backup and then unlinked on commit, and the run
reported success. Build them in an hts-tmp directory beside the file
instead, and remove it once the last slot sharing it is done.
Closes#774
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Cache repair deletes the old cache before a rename it never checks
Both zip-repair paths removed the damaged cache, moved repair.zip onto
its name without looking at the result, and announced a successful
recovery either way. A refused move left the entries under a name
nothing reads and no cache at all. Go through hts_rename_over(), which
keeps the destination when it cannot replace it, and report the failure
instead of claiming success.
Closes#786
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* A failed filecreate() on an unknown-length re-fetch commits the backup away
The re-fetch backup was dropped on commit without checking that the new
copy exists. filecreate() can fail after the backup was taken (EMFILE,
ENOSPC, EROFS), and back_finalize()'s incomplete-transfer guard is gated
on a known Content-Length, so a chunked response went straight to the
commit and lost both copies. Restore the backup instead when there is
nothing to commit to.
Closes#775
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* hts_rename_over() can lose its destination when the retried rename fails
The unlink-then-rename fallback leaves nothing in place of dst between
the unlink and the retry, so a retry that fails too loses it. Park dst
under a free scratch name instead, drop it once the move succeeded, and
put it back otherwise.
Closes#790
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Give the cache-repair test Windows coverage
The Windows job iterates a fixed glob, so a 106_engine-* test never ran
there; name it in the list. The move it exercises is the one whose
failure mode is Windows-specific.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Give the re-fetch backup test Windows coverage
The Windows job iterates a fixed glob, so a 108_engine-* test never ran
there. The rename semantics this drives are exactly what differs on the
CRT, which reserves EACCES for a locked source.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Review fixes: check the restore, probe in UTF-8, state the honest guarantee
The move back out of the parked name was unchecked, so a retry that
failed for a reason that still applied left dst absent with the content
orphaned under a name nothing reported. Check it, retry once, and name
the parked copy in the log; hts_rename_over() takes an httrackp for that.
The aside probe used fexist(), which is not UTF-8 and consults the ANSI
codepage on Windows while the renames beside it are wide. It also reads
a directory as a free name, so the park now skips a name whose rename
refuses rather than giving up on it.
The header claimed a failure leaves dst as it was, which the crash
window between the two renames does not give.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Reserve hts-tmp and hts-cache so no URL can name a temporary
Moving the temporaries into an hts-tmp directory only pushed the
collision one level down: hts-tmp is a path segment a URL can name, so a
site serving /d/hts-tmp/a.bin.bak still landed on the backup path for
/d/a.bin and was consumed. url_savename now escapes hts-tmp and
hts-cache as whole path components, the way it already escapes the DOS
device names, which also stops a URL from naming hts-cache/new.zip.
The three tests now drive the hostile shape rather than the pre-fix
name: a savename assertion on the escape, a crawl serving
/bakname/hts-tmp/a.bin.bak, and a selftest pinning the temporary's name.
Also: retry the backup once when another slot removed the shared
directory between structcheck and the rename; report a commit that had
to restore, so the caller does not cache the new validators against the
old body; drop the coded body on the too-long branch as its sibling
does; and scope the leftover scan to plants inside hts-tmp only.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* The aside fallback parked a directory that stood in the way
Windows refuses every rename onto an existing target, so the fallback is
production code there rather than the rare path it is on POSIX. A
directory at the destination was renamed aside like a file, the move
then succeeded, and UNLINK could not drop the parked directory, so the
call reported success where master had reported failure and left an
orphan behind. 101_local-update-stale-bak plants exactly that shape and
caught it on both Windows legs.
Park a regular file only. A directory in the way is refused, as before.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* A 304 during --update leaks the whole previous htsblk
back_wait() handles a 304 by replacing the response struct with the cache
entry, carrying only the socket and keep-alive members across via
back_connxfr(). The struct assignment drops every owned pointer the live
response still held without freeing any of them: the 8 KB header buffer on
every update, plus the two WARC header stashes when --warc-file is on. An
update over a 10k-page site drops roughly 80 MB in one run.
back_clear_entry() already knew how to tear those down, so the frees move into
a helper that both it and the 304 path call.
The new test runs the two-pass mini304 crawl with LeakSanitizer on, which the
sanitized CI job otherwise disables. The fresh first pass is the control: it
has no cache entry to read back and is clean either way. The update pass
reports 16 KB in 2 objects on master, one per unchanged URL, and nothing with
the fix.
Closes#782
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Trim the test header
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Review fixes: reuse deleteaddr(), and cover the WARC limb
back_free_response() was reimplementing deleteaddr(), which already frees adr
and headers and NULLs both; call it instead so the two cannot drift.
Test 114 never passed --warc-file, so the warc_free_request() limb ran with
both pointers NULL on every path it exercised and deleting it kept the test
green. A third pass turns the archive on, and it now fails with the 835 and 238
byte stashes when that call goes away.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Register the new leak test as an expected Windows skip
The Win32/x64 job pins the exact set of tests allowed to skip, so an
all-skipped suite cannot report green. 114_local-update-304-leak needs a
LeakSanitizer build and MSVC has no equivalent, so it skips there and tripped
the gate with fail=0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* A 304 revisit archives none of the exchange it stands for
The cache entry overwrites the whole htsblk in back_wait(), taking the
stashed request and response headers with it, so the revisit record gets
a synthesized "HTTP/1.1 200 OK" block and no request record. Both blocks
now move across the swap, but only for a real 304: the engine also forces
HTTP_NOT_MODIFIED by itself, and those have no 304 to carry. Since a 304
declares no Content-Type, warc_write_transaction() takes the resource type
from the caller so revisit CDXJ lines keep their mime.
cache_read_including_broken() returns the entry's response struct after
back_clear_entry() has freed it; adr, headers and location come back NULL
now, as the normal cache_readex() path already returns them.
Closes#826
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Do not expand an empty bodyargs array under macOS bash 3.2
Test 122 asks local-crawl.sh only for the revisit checks, so it is the
first caller to leave bodyargs empty. Bash before 4.4 calls that unbound
under set -u, which passes on Linux and fails the macOS leg.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Restore the opt argument to hts_rename_over call sites
Lost when the master merge took this file wholesale; #816 gave the function an
httrackp parameter and the three call sites here reverted to the old form.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* ProxyTrack's .ndx reader trusted an unparsed offset and the length of a path
An .ndx entry whose position field does not parse left "pos" untouched, so it
kept the offset of the entry before it, and the first entry read whatever the
stack held. The cache then resolves that URL to a different record, or to
nothing at all.
The base path and the .dat/.ndx filenames were built with strncat calls
bounded by the source path rather than by their 1024-byte destinations, so a
cache named on the command line with a long enough path overran them: ASan
reports a heap overflow on the filenames, UBSan an out-of-bounds index on the
base path right after. The zip reader carried its own copy of the same base
path code, reachable once the archive opens, and both now share one bounded
helper. A path that cannot fit is dropped rather than clipped, since a
truncated one names a different file.
Closes#825
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Skip the deep-path half where the filesystem will not hold the path
macOS caps a single path at 1024 bytes, so the 1200-byte tree the base-path
cases need cannot be created there and mkdir -p would fail the test outright.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* hts_newthread() leaves the pthread attributes object undestroyed when thread creation fails
hts_newthread() folds pthread_attr_init(), pthread_attr_setstacksize() and
pthread_create() into one short-circuit condition, and destroys the attributes
object only on the branch where all three succeeded. If setstacksize or create
fails (EAGAIN under thread or memory exhaustion being the realistic case), attr
stays initialised for good. glibc allocates nothing for a default-initialised
attr, so nothing actually leaks on Linux. This is a portability and hygiene fix
rather than a live bug.
The init has to come out of the chain instead of just gaining a destroy call:
the shared condition cannot tell which of the three failed, and destroying an
object pthread_attr_init() never initialised is undefined.
The imbalance is invisible to a leak checker, so the test counts it directly. A
small LD_PRELOAD shim refuses pthread_create() and records which attributes
objects nobody destroyed afterwards: 1 undestroyed on master, 0 with the fix.
Closes#772
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Review fixes: tighter comments, and only attr-bearing refusals count
Condense the two-line note above pthread_attr_init(), fix the "so that is can
be independent" typo on the line the diff moves, and count a refused spawn as
tested only when it carried an attributes object, so a NULL-attr create from
elsewhere cannot satisfy the guard on its own. Note why the shim's fixed slot
count and unlocked counters are enough.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Backward scans from strlen(s) - 1 walk off the front of the buffer
optinclude_file()'s config-line right trim and url_savename()'s collision
rename both start at the last character and decrement with no lower bound.
An all-space httrackrc line and an all-digit save name walk below their
stack buffers; the second is reachable from a crawl.
Both go through hts_rtrimlen(), which counts down from the end. The
Content-Disposition trim was a third copy and now shares it.
Closes#814
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* x + strlen(x) - 1 points before the buffer on an empty string
The pointer spelling of #770. All 27 occurrences go through
hts_lastcharptr(), which clamps to the terminating NUL, and the two
hand-written ternary guards from #729 and #767 fold into it.
Closes#781
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Cover the collision rename on Windows and from a followed link
116_engine-rtrim ran nowhere on Windows: that job iterates a fixed glob
that 116_engine-* does not match, and the rename is savename logic, which
behaves differently there. Add it by name.
The test only drove the rename through -g, which pins depth to 0 and so
needs both colliding URLs on the command line. Under -N "%n.%t" depth is
unrestricted, so one starting URL is enough and the crawled page picks
both names itself. Also strip CR before the empty-option check, which a
CRLF log would otherwise pass vacuously.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* size_t i = strlen(s) - 1: the (i > 0) guard cannot see the underflow
An unsigned index seeded with strlen(s) - 1 becomes SIZE_MAX on an empty
string, and SIZE_MAX passes the (i > 0) the sites carry. Seed from
hts_lastcharoffset() instead, and cover the spelling in the lastchar
self-test and its source guard.
Closes#821
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Drop the historical narration from hts_lastcharptr's contract
The declaration is the API surface, so it keeps the double-evaluation
warning a caller cannot derive from the signature; what the macro
replaced is git's job. Its grep exclusion in the test goes with it,
since that comment was the only htssafe.h line the scan matched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Catch the bare-reassignment spelling in the source guard
The declaration pattern cannot see punycode.c, where input_length is
declared a dozen lines above the subtraction, so reverting that fix alone
left the guard silent. Match an uncast assignment too; the (int) casts in
htscore.c and htstools.c are the safe form and stay out.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* ProxyTrack's .arc writer overflowed its 8192-byte header block
Converting a cache to .arc builds each entry's response headers with sprintf
into a fixed char headers[8192]. The one bound guarding the cached header
block compared its length against sizeof(headers) - strlen(headers) - 1, a
subtraction that wraps once the string reaches 8192 bytes, so the guard passed
precisely when it should have failed. The Location: write and the closing CRLF
had no bound at all. Every value comes back off a cache, and an .arc whose
entry carries enough unknown header lines is enough: ASan reports a 12 KB
stack write.
Each append now clips to the room left, holding back four bytes for the blank
line that ends the block and for the CRLF a clip landing mid-line still owes
its line. Clipping instead of aborting is deliberate: these are cache reads,
where a short record beats killing the conversion.
Closes#820
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Bracket the .arc header sweep around the measured cap
The fixed 7950..8250 window was 300 conversions picked to survive any change
to what the writer emits on its own. Measure that from the control entry
instead and sweep 32 bytes either side of the cap, which is the same coverage
in a fifth of a second.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Report a clipped .arc header block, and tighten the test around the real cap
Review follow-ups on the header-block bound. Truncating a cached header block
was silent, so a short record downstream had nothing pointing back at it; the
writer now names the URL on stderr, the way the readers already report a
corrupted entry. The test's own cap assertion allowed 8192 where the block
tops out at 8191, so a one-byte overflow would have passed it. runlen() had
been copied from the long-fields test and now lives in testlib.sh, and the
comment on the four reserved bytes says what depends on them.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Backward scans from strlen(s) - 1 walk off the front of the buffer
optinclude_file()'s config-line right trim and url_savename()'s collision
rename both start at the last character and decrement with no lower bound.
An all-space httrackrc line and an all-digit save name walk below their
stack buffers; the second is reachable from a crawl.
Both go through hts_rtrimlen(), which counts down from the end. The
Content-Disposition trim was a third copy and now shares it.
Closes#814
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Cover the collision rename on Windows and from a followed link
116_engine-rtrim ran nowhere on Windows: that job iterates a fixed glob
that 116_engine-* does not match, and the rename is savename logic, which
behaves differently there. Add it by name.
The test only drove the rename through -g, which pins depth to 0 and so
needs both colliding URLs on the command line. Under -N "%n.%t" depth is
unrestricted, so one starting URL is enough and the crawled page picks
both names itself. Also strip CR before the empty-option check, which a
CRLF log would otherwise pass vacuously.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* x + strlen(x) - 1 points before the buffer on an empty string (#819)
* x + strlen(x) - 1 points before the buffer on an empty string
The pointer spelling of #770. All 27 occurrences go through
hts_lastcharptr(), which clamps to the terminating NUL, and the two
hand-written ternary guards from #729 and #767 fold into it.
Closes#781
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Drop the historical narration from hts_lastcharptr's contract
The declaration is the API surface, so it keeps the double-evaluation
warning a caller cannot derive from the signature; what the macro
replaced is git's job. Its grep exclusion in the test goes with it,
since that comment was the only htssafe.h line the scan matched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Cache repair deletes the old cache before a rename it never checks
Both zip-repair paths removed the damaged cache, moved repair.zip onto
its name without looking at the result, and announced a successful
recovery either way. A refused move left the entries under a name
nothing reads and no cache at all. Go through hts_rename_over(), which
keeps the destination when it cannot replace it, and report the failure
instead of claiming success.
Closes#786
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* hts_rename_over() can lose its destination when the retried rename fails
The unlink-then-rename fallback leaves nothing in place of dst between
the unlink and the retry, so a retry that fails too loses it. Park dst
under a free scratch name instead, drop it once the move succeeded, and
put it back otherwise.
Closes#790
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Give the cache-repair test Windows coverage
The Windows job iterates a fixed glob, so a 106_engine-* test never ran
there; name it in the list. The move it exercises is the one whose
failure mode is Windows-specific.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Review fixes: check the restore, probe in UTF-8, state the honest guarantee
The move back out of the parked name was unchecked, so a retry that
failed for a reason that still applied left dst absent with the content
orphaned under a name nothing reported. Check it, retry once, and name
the parked copy in the log; hts_rename_over() takes an httrackp for that.
The aside probe used fexist(), which is not UTF-8 and consults the ANSI
codepage on Windows while the renames beside it are wide. It also reads
a directory as a free name, so the park now skips a name whose rename
refuses rather than giving up on it.
The header claimed a failure leaves dst as it was, which the crash
window between the two renames does not give.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* The aside fallback parked a directory that stood in the way
Windows refuses every rename onto an existing target, so the fallback is
production code there rather than the rare path it is on POSIX. A
directory at the destination was renamed aside like a file, the move
then succeeded, and UNLINK could not drop the parked directory, so the
call reported success where master had reported failure and left an
orphan behind. 101_local-update-stale-bak plants exactly that shape and
caught it on both Windows legs.
Park a regular file only. A directory in the way is refused, as before.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* An FTP --update resumed a complete mirror with REST and spliced the old file into the new body
FTP sent REST whenever the mirrored file merely existed, and on an --update
pass every previously mirrored file exists, so a complete copy was treated as
an interrupted download. The server resumed at its length and the mirror ended
up part old body, part new tail, at exactly the remote size, so nothing
downstream noticed. Resuming now follows the decision back_add() already makes
for HTTP, which only marks a copy partial when the cache does not hold it, and
r.size is seeded from the resume offset so a genuine resume is no longer
reported "FTP file incomplete".
Closes#798
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Format the two comment lines the earlier pass missed
git-clang-format only sees the diff present when it runs; the comments were
translated after it, so those lines never went through it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* An FTP resume spliced a changed remote when there was no cache entry
With no cache entry to tell a partial from a stale copy, back_add() took
range_req_size from the file size and the transfer resumed with REST, so a
remote that had changed since had its tail appended to the old head at exactly
the remote length.
FTP has no conditional retrieval, so the client has to decide. Resuming now
needs SIZE to report a remote strictly longer than the local copy and MDTM to
report it no newer than that copy. The mirror is stamped from MDTM as the HTTP
path is stamped from Last-Modified, so later passes compare two server-clock
times; a server that does not answer MDTM proves nothing and is re-fetched
whole.
Closes#823
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Clear r.lastmodified before an FTP attempt
run_launch_ftp() already resets msg, statuscode and size for a retry; without
lastmodified in that list a retry whose MDTM fails stamps the mirror with the
previous attempt's date.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Require the copy's date to match MDTM, not merely to lead it
"Remote no newer than the local copy" is one-sided, and it lets through exactly
the mirrors #823 is about: nothing stamped an FTP file before this branch, so
every existing one carries a client-clock mtime ahead of any MDTM, as does a
copy restored from an archive or moved without preserving times. The first pass
over such a tree resumed and spliced.
The date now has to match, which is what the HTTP path gets from the server via
If-Unmodified-Since. A partial written here is stamped from the same MDTM, so
the resume it exists for still compares equal.
Pass 6 plants a stale copy dated by the local clock over a larger changed
remote; it splices under the old comparison. Pass 5 now dates its copy as the
remote so the missing MDTM is the only thing refusing it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: fix TESTS list order and backslash after the union merge
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* --single-file dropped the fragment on an inlined reference
sf_resolve cuts a reference at '#' to find the mirrored file, and sf_inline
then replaced the whole reference with the data: URI it built, so the fragment
never came back. On an SVG that changes what renders: "sprite.svg#icon" selects
one element, and without the selector the browser draws the whole sheet.
Re-attach it to both replacements, the data: URI and the rebased path an
over-cap asset falls back to. The query stays dropped: it named the remote
resource, not the mirrored file.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Do not re-encode escapes the fragment already carried
The fragment is copied out of the document, so its '%' and '&' are the
document's own escapes; percent-encoding them again turns "#a&b" into a
lookup for "a&b". A mirrored path is the opposite case, a raw filesystem
name that has to be escaped, '#' included now that a fragment can follow it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Trim the test header and drop an unused top_srcdir
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Pin every member of the fragment escape set
Only ')' and the '%'/'&' pass-through were exercised, so dropping '"', a
quote, a paren, a backslash, '<'/'>', whitespace or the high-byte rule from
sf_append_escaped left the suite green. The two that would be a real
injection are the '"' that ends the attribute the rewriter re-quotes and the
whitespace that ends an unquoted url() token; the fixture now carries one
fragment per class, and each one dies to its own mutant.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Generate the test list instead of hand-maintaining it
Every branch that adds a test edited the same TESTS list in tests/Makefile.am,
so any two in flight collided there, and resolving that by taking one side drops
the other test silently -- nothing downstream notices a name that stopped being
listed. bootstrap now writes tests/tests-list.mk from tests/*.test and
Makefile.am includes it, so a test is registered by existing.
The generated file is untracked and shipped by make dist, the same arrangement
configure and the Makefile.in files already use, so tarball builds are unchanged.
Refs #844
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Track the test list rather than generating it at bootstrap
The first attempt had ./bootstrap generate tests/tests-list.mk, which broke
every build: CI regenerates with autoreconf -fi and Debian with dh_autoreconf,
neither of which runs bootstrap, so automake aborted on the missing include.
Track the fragment instead. One "TESTS += name" per line with no continuation,
and merge=union scoped to that file, so two branches adding a test union cleanly
instead of colliding -- the shape matters, since unioning a backslash-continued
list drops an entry silently, which is why #843 was closed.
EXTRA_DIST does not need the fragment; automake already ships an included file
via am__DIST_COMMON.
Refs #844
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Correct the comment that still described the generated design
tests/Makefile.am claimed the list was generated and that a test registers by
existing. Neither is true after the pivot to a tracked file, and a contributor
who believed it would add a .test and have it silently skipped.
Also record the one hazard union cannot avoid: a deletion merged alongside
another branch's append is restored, so a removal wants a commit of its own.
Refs #844
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Reading an .ndx cache with full-length URL fields reads past the buffer
binput() returned count + 1 unconditionally, assuming it had consumed a
separator. When it stopped on the buffer's terminating NUL, or because the
destination filled, that byte was not a separator and the caller's cursor
moved one past it. In PT_LoadCache__Old the cursor then leaves the heap
buffer holding the .ndx and the next binput() call reads out of bounds; the
loop's a < (use + ndxSize) guard cannot see it, since the drift happens
after the check.
Step over the byte only when it really is a newline. cache_brstr() carried
a caller-side patch for the same drift and had to follow, or its adr[off-1]
probe would read before the buffer once binput() can return 0.
Closes#793
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Broaden test 117 and make its bookkeeping fail closed
The trigger is the walk reaching the buffer's terminating NUL, which any
.ndx not ending in a newline does; full-length fields are one route to it,
not the only one. Add a field that stops short of every bound, and a pair
straddling cache_brstr's own 256-byte field, the second destination the
diff touches.
items() returned 0 whenever the item-count line was missing, so a crawl
that printed nothing passed the zero-key assertion. Signals now clean up on
their own trap line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Rename test 117's zero-byte fixture off the Windows NUL device
Windows resolves NUL.<anything> to the null device, so nul.ndx never
existed there and proxytrack reported an unloadable index. The suite's
Win32 leg caught it once items() stopped treating a missing count as zero.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* A 304 during --update leaks the whole previous htsblk
back_wait() handles a 304 by replacing the response struct with the cache
entry, carrying only the socket and keep-alive members across via
back_connxfr(). The struct assignment drops every owned pointer the live
response still held without freeing any of them: the 8 KB header buffer on
every update, plus the two WARC header stashes when --warc-file is on. An
update over a 10k-page site drops roughly 80 MB in one run.
back_clear_entry() already knew how to tear those down, so the frees move into
a helper that both it and the 304 path call.
The new test runs the two-pass mini304 crawl with LeakSanitizer on, which the
sanitized CI job otherwise disables. The fresh first pass is the control: it
has no cache entry to read back and is clean either way. The update pass
reports 16 KB in 2 objects on master, one per unchanged URL, and nothing with
the fix.
Closes#782
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Trim the test header
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Review fixes: reuse deleteaddr(), and cover the WARC limb
back_free_response() was reimplementing deleteaddr(), which already frees adr
and headers and NULLs both; call it instead so the two cannot drift.
Test 114 never passed --warc-file, so the warc_free_request() limb ran with
both pointers NULL on every path it exercised and deleting it kept the test
green. A third pass turns the archive on, and it now fails with the 835 and 238
byte stashes when that call goes away.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Register the new leak test as an expected Windows skip
The Win32/x64 job pins the exact set of tests allowed to skip, so an
all-skipped suite cannot report green. 114_local-update-304-leak needs a
LeakSanitizer build and MSVC has no equivalent, so it skips there and tripped
the gate with fail=0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* hts_rename_over() can lose its destination when the retried rename fails
The unlink-then-rename fallback leaves nothing in place of dst between
the unlink and the retry, so a retry that fails too loses it. Park dst
under a free scratch name instead, drop it once the move succeeded, and
put it back otherwise.
Closes#790
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Review fixes: check the restore, probe in UTF-8, state the honest guarantee
The move back out of the parked name was unchecked, so a retry that
failed for a reason that still applied left dst absent with the content
orphaned under a name nothing reported. Check it, retry once, and name
the parked copy in the log; hts_rename_over() takes an httrackp for that.
The aside probe used fexist(), which is not UTF-8 and consults the ANSI
codepage on Windows while the renames beside it are wide. It also reads
a directory as a free name, so the park now skips a name whose rename
refuses rather than giving up on it.
The header claimed a failure leaves dst as it was, which the crash
window between the two renames does not give.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* The aside fallback parked a directory that stood in the way
Windows refuses every rename onto an existing target, so the fallback is
production code there rather than the rare path it is on POSIX. A
directory at the destination was renamed aside like a file, the move
then succeeded, and UNLINK could not drop the parked directory, so the
call reported success where master had reported failure and left an
orphan behind. 101_local-update-stale-bak plants exactly that shape and
caught it on both Windows legs.
Park a regular file only. A directory in the way is refused, as before.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The FTP worker writes url_sav itself, so its slot carries a size but no
in-memory body. Serializing that slot to the on-disk ready table stores no
body, and the read-back took the size from what it stored, leaving zero: the
link writer then saw an empty response and created a 0-byte file over the
bytes already on disk, while the engine logged the transfer as a success.
Test 110 mirrors twelve files at -c8, which is what makes a ready slot wait
long enough to be swapped, and -#test=backswap covers the round-trip directly.
Closes#797
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* PT_GetTime copied gmtime's shared static instead of a reentrant breakdown
On _WIN32 the success path took gmtime()'s pointer and dereferenced it after
the fact, so a concurrent conversion on another thread could change the
breakdown under it. The POSIX branch was already reentrant via gmtime_r, and
the same #ifdef pair had been copy-pasted into hts_now_iso8601() and the WARC
auto-name; fold all of them onto one hts_gmtime() helper, and give ProxyTrack's
WebDAV listing the same treatment, since it read the static's fields well past
the call.
Windows uses Microsoft's gmtime_s (destination first, errno_t return), not the
C11 Annex K function of the same name.
Covered by a new "gmtime" engine self-test: a reference table checks the
breakdown itself, which is what catches a swapped-argument call on the MSVC
leg, and eight threads hammering the helper catch a return to the shared
static (16k of 400k conversions corrupt with that mutant in place).
Closes#794
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Give the gmtime self-test teeth on the failure path and off UTC
Three holes the test-design audit found by running mutants rather than
reading the diff.
A helper that discarded gmtime_r's NULL and always claimed success passed
every phase, yet that boolean is the only failure signal hts_now_iso8601,
warc_open and PT_GetTime have; all three would have formatted an
uninitialised struct tm. A forced-failure row now converts INT64_MAX, gated
on a 64-bit time_t.
The localtime_r mutant only died on a non-UTC box. CI runners are UTC, where
localtime_r and gmtime_r agree on every reference row, so the test exports
TZ=XXX5.
The "first result survives the second call" phase could not fail: both
buffers are caller-owned stack storage no implementation writing through
tmbuf could disturb. Removed rather than left reading as coverage.
expect_ok() was a third byte-identical copy; it moves to tests/testlib.sh
with the two existing callers.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Rename the self-test's out-of-range time_t off the "far" keyword
WinDef.h defines "far" away to nothing, so the declaration lost its variable
and MSVC rejected the file. Both Windows legs caught it; the POSIX builds
never see the macro.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
(void) does not suppress glibc's warn_unused_result, so gcc warned on
the read-back in st_logcallback. Treat a failed read as the test failure
it is instead of asserting against an empty buffer.
Closes#812
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Match captured output with a here-string, not a pipe into grep -q
grep -q exits on the first match, so whatever the producer still had to
write takes SIGPIPE; under pipefail that becomes the pipeline's status and
an assertion that held reports failure. bash issues one write() per line,
so any match that is not on the last line is exposed.
Converts every test assertion whose producer is a shell builtin or shell
function, including two pipelines used as an if condition where the SIGPIPE
silently flips the branch. Generalizes the AGENTS.md bullet, which only
covered the "&& fail" spelling.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Convert the remaining pipe-into-grep -q sites
The self-test drivers survive only because the matched line is the last
thing httrack prints; one added line of output turns a pass into 141.
Nothing in tests/ pipes into grep -q now.
The two zlib drivers claimed the harness might run them under a POSIX
/bin/sh: it does not. configure resolves $(BASH) to bash, test-timeout.sh
execs it, and 01_zlib-warc-wacz.test already uses "set -o pipefail" (which
dash lacks) on the macOS leg. Kept the half that is true, BSD tool flags.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* The -o help text promises a generated error page the engine never builds
`-o` only decides whether the error page the server sent survives: `store_errpage`
keeps `r.adr` alive so the normal save path writes it, and the `-o0` arm frees it.
Nothing anywhere builds a stand-in body. The one block that would have was dead
since the 3.20.2 import and was removed in #783.
Reword the help line, the man page and fcguide's two `-o` prose blocks to say the
server's error page is saved rather than generated, and extend 23_local-errpage
so the mirrored 404 has to carry the server's own body.
Closes#787
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Condense the -o1 control comment to one line
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* An FTP --update resumed a complete mirror with REST and spliced the old file into the new body
FTP sent REST whenever the mirrored file merely existed, and on an --update
pass every previously mirrored file exists, so a complete copy was treated as
an interrupted download. The server resumed at its length and the mirror ended
up part old body, part new tail, at exactly the remote size, so nothing
downstream noticed. Resuming now follows the decision back_add() already makes
for HTTP, which only marks a copy partial when the cache does not hold it, and
r.size is seeded from the resume offset so a genuine resume is no longer
reported "FTP file incomplete".
Closes#798
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Format the two comment lines the earlier pass missed
git-clang-format only sees the diff present when it runs; the comments were
translated after it, so those lines never went through it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The changed-lines clang-format job resolved its base as origin/<base_ref>,
which is master's tip when the job runs. Once master gains a C commit while
a PR is open, the comparison also picks up the reverse of that commit and
the job fails on code the PR never touched.
Use git merge-base instead, and fail loudly if there is none rather than
falling back to a whole-tree comparison.
Closes#800
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
hts_log_vprint() makes a va_copy and then throws it away: the callback is handed
the original args, and vfprintf() writes the log file from that same,
already-consumed list. On x86_64 a va_list is a one-element array, so the callee
moves the caller's cursor; the second traversal reads past the register save
area, and a %s yields a junk pointer that vfprintf() dereferences.
Only an embedder that installs a callback is affected, so the CLI never sees it.
HTTrack Android does: --sitemap is the first option whose LOG_NOTICE lines carry
arguments, and ticking its checkbox segfaults the crawl thread. The same crawl is
clean under the CLI built with ASan+UBSan.
Pass the copy to the callback. -#test=logcallback sends one line with a %d and a
%s through both sinks and compares them; on the unfixed code the log file gets
"0 " and the test fails. It also logs a second line below opt->debug with no
opt->log, pinning that the callback fires above the level filter.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Under `set -e` a failing command in an EXIT trap becomes the script's exit status, so a hiccup while tearing down fixtures fails a test whose assertions all passed. That is what turned `57_local-proxy-connect.test` red on the Windows x64 leg of #765: five OK lines, no FAIL, exit 1. Every EXIT trap in the suite now runs teardown with errexit off, and the signal traps keep their own `trap` line, since sharing `set +e` with HUP/INT/QUIT/PIPE/TERM would leave errexit off for the rest of a signalled run and let a torn-down test still report success.
A `|| true` on the `rm` would have been smaller, but it throws away the only diagnostic, and the evidence does not say which teardown command failed: a blocked `rm -rf` on the Windows runner exits 1 and prints "Device or resource busy", while the log shows exit 1 and nothing at all. The sharing violation in the issue is the plausible mechanism rather than a confirmed one, so whatever it really is now prints its own error.
`99_teardown-status.test` pins the semantics both ways and scans the suite so a new test cannot reintroduce the shape, the leaky combined trap included. The `return 0` that three `cleanup()` bodies ended with never protected anything, since errexit fires at the failing command before it is reached.
Closes#773
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A test that wedges runs until CI cancels the step, and a cancelled step keeps neither its log nor the artifacts its `if: always()` uploads would have produced. That is why nobody has ever been able to say which Windows test hangs, across 19 dead jobs in the last day alone (#795).
Each test now runs under a wall-clock budget at the automake harness level, so an overrun names the test, dumps the surviving process tree and an engine stack, and exits 124. The step then fails rather than being cancelled, which is what keeps the log. Every POSIX `make check` leg gets this; the Windows leg runs its own serial loop and now calls the same wrapper. The budget is 600s, the value the Windows leg already used: it has to clear the 540s a three-pass crawl may legitimately take under `local-crawl.sh`'s own 180s-per-pass watchdogs, against a slowest healthy test that actually measures 39s.
Two things sit on top of the per-test bound. The Windows suite gives up at 25 minutes so it fails on its own terms well before the 45-minute step timeout, and it sweeps leaked engine processes between tests, naming whichever test left them. An orphaned `httrack.exe` starving the runner is the leading theory for the hang, and that sweep is what would confirm it.
This also fixes an unbounded `wait` after `kill_tree` in testlib.sh. When the kill failed to reap, which is exactly the native-Windows case those watchdogs exist for, the watchdog blocked forever and never printed the timeout it was about to report.
Stacks differ by platform, and each branch says which one it took, because a dump that silently produces nothing reads as coverage. Linux sends SIGABRT and lets httrack's own crash handler symbolize itself, verified against a real wedged crawl where it named `back_wait` at htsback.c:2710. macOS cannot do that, since htsbacktrace.c gates the handler on `__linux`, so it uses `sample(1)`. Windows uses `cdb` from the SDK, and that is the one path I could not exercise from here; it probes for the binary, reports when it is absent, and is bounded so a debugger that wedges cannot become the new hang.
Does not fix#795, only makes it diagnosable.
The between-test sweep costs the Windows step about two minutes (506s before, 619s on this run). That buys naming whichever test leaks, which is the whole lead on #795; it can be narrowed or dropped once the leak is found.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A 304 revisit record carried a `WARC-Profile` and nothing else naming what it stood for. Neither replay engine reads that field: pywb resolves in `_load_different_url_payload` on `WARC-Refers-To-Target-URI` and raises `ArchiveLoadFailed` without it, wabac.js reads `warcRefersToTargetURI` and otherwise answers Not Found. So the records were conformant under WARC 1.1 6.7.3, which only recommends the field, and unreplayable in both engines that matter.
For a server-not-modified revisit the referred-to URI is the record's own target URI, so it is free to emit.
No `WARC-Refers-To-Date` alongside it. The cache persists no capture timestamp: the field list in `cache_add()` ends at `Last-Modified`, which is a document property, and the zip entry's own date is set from that same value. Emitting it would assert that a record exists with that `WARC-Date`, which is false and would misdirect pywb's `closest=` lookup. Both engines already handle the field being absent, pywb by falling back to the CDX timestamp and wabac.js to the revisit's own. A real date needs a capture time in the cache, which is worth doing when the segment work lands and a previous index is being read anyway.
The validator now requires every revisit to name a capture, and requires a server-not-modified one to name its own URI. Test 73 already drives it over an archive of revisits; against the pre-fix engine it fails there.
Stacked on #788. Without it the new header line is what pushes a 995 to 1004 byte URL past `wbuf_printf`'s old 1024-byte buffer, and `warc_emit` then drops the record whole; the review caught that before this was pushed. The merge is in the branch so the pair is what got tested.
Closes#778
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every WARC header line is built with `wbuf_printf()`, which formatted into a 1024-byte stack buffer and returned `-1` instead of growing. `"WARC-Target-URI: %s\r\n"` costs 19 fixed bytes, so the line failed once a URL reached 1005 bytes, the `-1` reached the `goto done` in `warc_emit()`, and the record was abandoned. The page still got mirrored, the crawl still exited 0, and nothing was logged, so the only symptom was a URL missing from the archive.
`wbuf` reallocs already, so oversized output now formats straight into it after a `wbuf_reserve()`, which is the growth half of `wbuf_add()` split out. Every field benefits, not only the two carrying URLs. The second pass is bounded against what was reserved rather than trusted: advancing `len` by a return value larger than the reservation would push `len` past `cap` and corrupt the bounds check of every later append.
`-#test=warc-longurl` sweeps 100 to 9000 bytes across the boundary. Against the old formatter it fails at exactly 1005 and up, with 1003 and 1004 passing. Each record carries a distinct payload because identical ones dedupe into revisits, which would otherwise hide the response records the test counts.
One caveat on the sanitizer evidence, since it is easy to over-read. The buffer grows by doubling, so a small off-by-one lands in allocation slack where ASan cannot see it; that is why the second-pass bound is a logic check rather than something left to the sanitizer. The ASan+UBSan run over the sweep is clean, but only after planting a deliberate overflow in `wbuf_reserve` to confirm the probe actually fires. It did not, at first: libtool silently drops `-fsanitize` from the shared-library link, which produced no binary at all and a "clean" result that meant nothing.
Worth knowing for the segment work: `warc_emit()` returning `-1` also sets `w->failed`, which suppresses the archive swap added in #777. Before this fix a single over-long URL would make an `--update` pass throw away its whole archive and keep the previous one.
Closes#785
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An FTP re-fetch called `filecreate()` on the mirrored file before a byte of the transfer had arrived, so a read error, a timeout or a short body destroyed the previous copy. HTTP has moved the good copy aside to a `.bak` since the #77 follow-up and puts it back when the transfer fails; FTP never took that backup. Rather than give it a second copy of the idiom, the backup moves out of `back_wait()`'s direct-to-disk block into `back_refetch_backup()`, which the FTP transfer now calls too. `back_finalize()` already restores it, so FTP inherits that. The REST resume branch appends and needs no backup.
That alone was not enough. `back_cleanup_background()` swaps a ready slot to the on-disk table through `back_clear_entry()`, which unlinks `back->tmpfile`, and a failed re-fetch sits at `STATUS_READY` unfinalized until the parser picks it up. The new `.bak` was being deleted in that window before `back_finalize()` could restore it, about half the time on a loaded box. `slot_can_be_cached_on_disk()` now refuses a slot that still owns a temporary, which closes the same window on the HTTP `.bak` and on the content-coding spool. The crawl-level race needs concurrency the suite cannot pin down, so `-#test=backswap` covers the predicate directly.
The suite had no FTP server. `tests/ftp-server.py` is a minimal one (PASV, SIZE, REST, RETR, LIST) with a mode file the test rewrites between passes, so a path can start failing without the port moving and taking the mirror directory name with it. Test 102 mirrors three files, then re-fetches one cut short, one served empty and one healthy: the first two must come back byte identical and be reported `unchanged`, the third replaced. It runs at `-c1` because a parallel FTP crawl loses whole transfers to #797, which is older and unrelated and would flake it on a loaded runner. #798 came out of the same work: an FTP `--update` sends `REST` over a complete mirror and splices the old file into the new body.
Closes#771
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The unlink-then-rename fallback in `hts_rename_over()` is there because Windows' rename() refuses an existing target, but it fired on any failed rename, ENOENT on the source included. A caller moving a temp file it never managed to write lost the destination and got `HTS_FALSE` back, which reads as "nothing happened". #754 unified four hand-rolled copies into this helper, so every call site inherited it.
The unlink now runs only for EEXIST, and only with a source that exists. EEXIST is the value that matters: the CRT maps ERROR_ALREADY_EXISTS there and keeps EACCES for a source another process holds, so accepting EACCES as well would have deleted the destination for a failure it had no part in, and the retry would then fail with the destination already gone. The source check is belt and braces for a CRT that reports neither. `hts_rename_utf8()` preserves errno across its free() calls now, since the gate reads it.
The existing call sites all derive their source's existence from a create that had to succeed first, so the bug was latent rather than live, but three of those destinations are user data: a mirrored file, a finished .wacz, and a rewritten page nothing will re-fetch. What is left of the window after this is #790.
The selftest probes what rename() does to an existing target and asserts against the regime it finds, so it runs three ways on Linux: native, then under an LD_PRELOAD rename() with Windows' shape, then under one reporting a locked source. The harness pins the expected regime per platform, so no leg can pass having exercised the other half. Seven mutants of the gate were each confirmed to fail it.
Closes#779
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three bugs have already come out of the same one-liner: `x[strlen(x) - 1]` indexes one byte before the buffer when the string turns out to be empty (#729, #730, #768). Rather than wait for a fourth, this replaces the idiom everywhere with helpers that cannot underflow.
All 48 occurrences were reclassified by tracing each guard to where it actually lives instead of to the two lines above the index. 44 were already guarded, often a dozen lines up or in the caller, and one sits inside a commented-out block. The rest have no guard at the site, and two of those are reachable from a crawl.
`htscore.c:2194` is the one that matters: a one-byte stack out-of-bounds write in the end-of-mirror purge, confirmed under ASan and UBSan. `linput()` reads `old.lst` in 999-byte chunks, so a 1000-byte line comes back as 999 bytes plus a one-byte tail; for that tail `line + 1` is empty, and without `-O` so is `path_html`, which leaves `file` empty when the index runs. The crawled site controls the save-path length and therefore the line length. It needs a second run over the same project, which is what a re-crawl does. As controls, a 900-byte path and the same path with `-O` both stay clean.
`htsparse.c:2009` is a one-byte overread in the HTML parser, reachable from `<a href=" ">`, where the guard sits on the wrong side of the `&&`. The write on the next line is accidentally safe for the same reason. #768 itself turned out not to be reachable from a crawl, only through the `-#test=savename` hook, so I would not call that one a security fix.
The helpers live in `htssafe.h` beside the other bounded string operations, so no file needed a new include. Guards doing more than an emptiness check are kept, since folding those in would append a separator to an empty string.
`-#test=lastchar` checks the helpers against a poisoned neighbouring byte, including the `/`-before-the-buffer case from #768, and putting the missing length check back makes it fail. It also greps the source, so reverting any converted site fails instead of leaving the self-test green. `97_local-purge-longpath` drives the purge site end to end; on unfixed code the sanitized CI leg aborts at `htscore.c:2194`.
Built from identical source paths, 61 of 72 objects are byte-identical, one differs only in debug info, and the remaining 10 are the files edited.
The pointer spelling `x + strlen(x) - 1` has 25 occurrences and is filed separately as #781.
Closes#770Closes#768
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The stand-in body `httpmirror()` builds for an error page has been unreachable since the 3.20.2 import: both guards say the opposite of their comments, so the block runs only when there is no save name and the URL *is* `/robots.txt`. `create_html_warning` is never assigned, so the HTML arm is dead outright. The GIF arm fires only if the user maps `.txt` to `image/gif` with `--assume`, and it then swaps `r.adr` for a 1070-byte buffer while leaving `r.size` at the error body's length, so the `robots_parse()` call below over-reads the heap.
Repair is not the one-character fix the inverted guards suggest. `filesave()` below writes `r.size` bytes, so uninverting the guards without also setting `r.size` moves the same over-read into the user's mirror. A correct repair would then replace every server error body with HTTrack's 2003 template by default, which is what `23_local-errpage.test` asserts is kept. Nobody has asked for the placeholder in twenty years, and #17 asked for the opposite.
`--errpage` itself is untouched: `store_errpage` keeps the server's error body and the normal save writes it. A byte-level differential against a master build over five error shapes under both `-o1` and `-o0` gives identical mirrors, and rebuilt objects differ only in `htscore.o`. The new test drives the one input that reached the block, and fails on master in a plain build rather than only under the sanitizers. Separately, the `-o` help text promises a generated page the engine has never produced; that predates this change and is filed as #787.
Closes#769
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`back_finalize()` moves an existing file aside to `<file>.bak` before truncating it, so an aborted re-fetch can put the previous copy back. That rename was a bare `RENAME`, and Windows refuses to rename onto an existing target, so a `.bak` outliving a killed run made every later re-fetch of that URL fail the rename, take the `tmpfile = NULL` branch with nothing logged, and truncate the live file with no backup at all. The #77 guard was off for that file, silently and for good. `hts_rename_over()` (#754) unlinks and retries, which fixes it.
Worth a look, because clobbering a stale `.bak` is not free. A run killed between the rename-aside and the finalize leaves the previous complete copy in `.bak` and a partial in the live file, and the next re-fetch now overwrites the good one. It still looks like the right trade: POSIX `rename()` has behaved exactly this way since #77 landed, so the alternative is leaving Windows with a guard that stays dead until someone deletes the file by hand, and a leftover `.bak` is engine garbage that nothing advertises or ever restores. The clobber is logged now, so it is at least visible.
Any failure to create the backup is logged too, instead of quietly disabling the safety net. `htscache.c`'s static `hts_rename()` wrapper and the hand-rolled `old.zip` unlink at its only call site go the same way, which removes a copy of the unlink-then-rename idiom rather than adding a fifth.
Test 101 plants both kinds of leftover before the update pass: a stale file, which has to be clobbered so the `-M` abort can still restore the pass-1 copy, and a directory, which cannot be clobbered and has to be reported instead. Only the Windows leg arms the first half, since POSIX clobbers on its own; the second is armed everywhere. Test 37 gains a live update pass so the dead pass rotates onto an existing `old.zip`, which nothing covered.
Two older bugs on the same path came out of the review and are filed separately: #774 (the `.bak` name collides with a mirrored file of the same name, reproduced) and #775 (a failed `filecreate()` on a chunked re-fetch commits the backup away).
Closes#758
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* warc: keep the previous archive when a pass has no bodies to replace it
A second crawl into the same output reopened the WARC with "wb" and
truncated it. On a cache-served pass nearly every URL comes back 304, so
the new file held revisit records whose payloads had just been deleted,
and the regenerated WACZ came out with zero page rows: a package that
replays nothing, in place of one that replayed fine.
The writer now builds into a sibling .tmp whenever an archive is already
there, and only swaps it in at close if the result can stand on its own.
A pass that only revisited URLs it did not re-download keeps the previous
.warc.gz, .cdx and .wacz untouched and says so.
What --update should ultimately mean for WARC output is still open; this
only stops the silent destruction in the meantime.
Closes#759
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* warc: guard the segment swap and cover the rotated archive
A run that lost a record or a segment must not replace a whole archive,
and hts_rename_over unlinks its destination when the source is missing,
so every segment has to be on disk before the first rename.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* warc: force rotation in the segment test and guard its vacuity
--warc-max-size 2000 never rotated under the harness's --robots=0, so the
segment test was checking a single-file archive; a mutant that renamed
only segment 0 passed it. 600 rotates, and --archive-min-files fails the
test if a shrinking crawl ever stops producing the segments it checks.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The MD5 comparison the comment promises left the tree in two steps: #467
extracted the function out of htsparse.c's HT_ADD_END macro without the skip
branch, and #512 removed the //[HTML-MD5]// cache entry it read. Drop the
parenthetical; the write is unconditional.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* A failed re-fetch overwrote the mirrored file with the aborted read's debris
A transfer that dies before a complete response has no body, yet the save path
still consulted r.adr, which at that point holds whatever the aborted header
read left behind: raw status-line bytes, or an empty buffer that truncated the
file to zero on macOS. Require a successful transfer, as the empty-body half of
the condition already did.
Closes#748
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: assert the failed re-fetch keeps its bytes on every platform
Test 93 filtered reset.bin out of its bucket lists because a connection killed
before the status line surfaced differently per platform. It no longer does, so
assert the resource like any other: unchanged, with the bytes pass 1 mirrored.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: make the re-fetch assertions catch a crawl that never re-fetched
The new test passed with no second pass at all, so a regression that stopped
re-fetching would have looked green. Assert the failure the fixture provokes,
give stay.bin a fresh pass-2 body so a fix that stopped overwriting anything
fails, and compare reset.bin by checksum rather than by length.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* --purge-old deleted a file whose re-fetch never got a response
A transfer that dies on the wire leaves the previously mirrored copy in place,
but back_finalize() returned without noting it, so the URL fell out of new.lst
and the end-of-update purge treated the file like a page that had vanished from
the site. Note the surviving copy, the way the incomplete-transfer branch above
already does, for the connection-level failures htsparse.c retries on. A
deliberate skip (too big, MIME-excluded, cancelled) keeps its current fate.
Closes#746
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Keep the copy for every failure, not just the five retryable codes
A malformed status line, an oversized declared length or a mid-flight abort all
land on STATUSCODE_INVALID, outside the retryable set, and the purge still ate
the file. Invert the test: any failure keeps the copy except the codes that
mean the engine passed the resource over on purpose.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: assert the retry exhaustion, not the per-platform failure message
The message a cut connection produces depends on whether any bytes arrived, so
matching it would fail on a runner that sees none.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* A failed re-fetch overwrote the mirrored file with the aborted read's debris
A transfer that dies before a complete response has no body, yet the save path
still consulted r.adr, which at that point holds whatever the aborted header
read left behind: raw status-line bytes, or an empty buffer that truncated the
file to zero on macOS. Require a successful transfer, as the empty-body half of
the condition already did.
Closes#748
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: assert the failed re-fetch keeps its bytes on every platform
Test 93 filtered reset.bin out of its bucket lists because a connection killed
before the status line surfaced differently per platform. It no longer does, so
assert the resource like any other: unchanged, with the bytes pass 1 mirrored.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: make the re-fetch assertions catch a crawl that never re-fetched
The new test passed with no second pass at all, so a regression that stopped
re-fetching would have looked green. Assert the failure the fixture provokes,
give stay.bin a fresh pass-2 body so a fix that stopped overwriting anything
fails, and compare reset.bin by checksum rather than by length.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: assert the retry exhaustion, not the per-platform failure message
The message a cut connection produces depends on whether any bytes arrived, so
matching it would fail on a runner that sees none.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
htsthread_wait_n(background_threads - 1) subtracts one more than the count of
threads that must not be joined. Without --ppid that count is zero, so the wait
asks for a negative number of outstanding threads and the counter never gets
there; with --ppid it waits on the pinger, which by design never returns.
Wait for background_threads instead, which is what the sibling call inside
back_launch_cmd() already passes.
Closes#753
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* htsAddLink walks back from an empty codebase, one byte before the buffer
Same idiom as the lienrelatif() underflow fixed in #729: the trim that walks
back to the last '/' starts at codebase + strlen(codebase) - 1, which is
codebase - 1 when the string is empty, and the loop dereferences it before
a > codebase stops it.
Unlike #729 there is no reachable empty value. codebase is copied from a
recorded link's fil, and no hts_record_link() call site can supply an empty
one: every fil is either a literal seed or an ident_url_absolute() /
ident_url_relatif() success return, and fil_simplifie() restores "/" or "./"
rather than leaving a path empty. The guard goes in anyway, and -#test=addlink
drives the walk directly: under the sanitize job's ASan+UBSan build the test
fails on the unfixed walk and passes with the guard. Its second case pins the
ordinary trim, which the guard leaves alone.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* review: add the case that actually notices the codebase trim
For an ordinary relative link ident_url_relatif() re-derives the directory
from the path it is handed, so deleting the trim outright left the two
existing cases green. A query-only link ("?x=1") copies that path whole, and
does catch it.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* structcheck() builds its rename target with an unbounded sprintf
Both structcheck() and structcheck_utf8() move a regular file sitting where a
directory belongs, and build the "<name>.txt" target with a raw sprintf into a
2048-byte buffer. It stays in bounds only because of a strlen(path) >
HTS_URLMAXSIZE guard dozens of lines above, which nothing at the write site
mentions. Route both through sprintfbuff() and fail with ENAMETOOLONG, so the
bound is local.
Armed the probe: with that distant guard patched out, a 2045-byte path makes
the old sprintf write 2049 bytes into tmpbuf[2048] under ASan; the same build
with sprintfbuff() returns -1 and reports nothing.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* review: tighten the structcheck self-test and its comments
The path builder could end a path with a bare separator when the base
directory's length hit the wrong residue, so the test aborted on a long
$TMPDIR. The refusal case also asserted on a component structcheck never
creates; assert on the outermost one instead.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: drop the max-length rename case, macOS PATH_MAX is 1024
The path the guard admits (HTS_URLMAXSIZE) plus ".txt" is longer than macOS
accepts, so fopen() failed there. The case could not tell a fixed build from an
unfixed one anyway; what is left covers the guard and the rename on both entry
points.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
configure has rejected --without-zlib since #750, and htsglobal.h
#errors on a forced HTS_USEZLIB=0, so the macro can only ever be 1.
Collapse every #if/#ifdef HTS_USEZLIB guard (htsweb.c, htsname.c,
htslib.c, htszlib.c, htsselftest.c, htscodec.c) to its always-taken
branch; htsweb.c's guard used #ifdef where every other site used #if,
an inconsistency that no longer matters once the guard is gone.
htscodec.c's #else arms were a genuine zlib-free content-coding path
(Accept-Encoding: identity, hts_codec_unpack returning -1), not stubs.
Removed for consistency with the other nine guards: the cache
(htscache.c, proxy/store.c) already reaches minizip unconditionally
from ~60 call sites with no null backend, so a zlib-free build is not
actually reachable today regardless of this file.
No new test: this is dead-code removal with no behavior change.
Verified by differential build against master: htscodec.o and
htszlib.o are byte-identical; the other touched objects differ only
in __LINE__ immediates shifted by the removed guard lines (plus one
cosmetic objdump label-annotation artifact each in htsselftest.o and
htsweb.o, from string-literal pool reordering). Exported symbols in
libhttrack.so are unchanged. make check: 166/166 (158 pass, 8 skip).
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
htsselftest.c and tests/Makefile.am were held by #718 while #747 was fixed, so
the thread-counting fix went in without a test. -#test=threadwait covers it
from both sides: a wait placed right after a spawn joins that thread, and
wait_n(n) leaves n running rather than draining them.
One spawn per round is what makes it bite. A batch gives the earlier threads
time to raise the counter themselves, which is why an eight-thread version
passed on the unfixed engine; one thread per round failed 10 runs out of 10.
The changes-race self-test can now drop the counter it kept because
htsthread_wait() could not be trusted to join.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* One unlink-then-rename helper for the three copies
replace_file() (htsback.c), wacz_rename_over() (htswarc.c) and the inline
block in singlefile_rewrite_file() each worked around Windows' non-clobbering
rename, with two opposite return conventions and only one of the three
converting path separators. Copying the wrong one gets you an inverted success
test.
hts_rename_over() replaces all four call sites: hts_boolean return, fconv() on
both paths. It lands in htsname.c rather than the htstools.c the issue named,
to stay clear of an in-flight PR over that file.
The wacz test now runs a cacheless second pass over a poisoned copy of the
package the first pass wrote, so repackaging has to clobber an existing file.
43 and 74 also assert the failure warnings stay out of the log, which is all
an inverted return leaves behind at those two sites.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
* rename helper: move hts_rename_over to htstools.c
Issue #726 asked for htstools.c; it went to htsname.c only because #718 held
the file at the time. Kept internal, not HTSEXT_API: htscore.h pulls in both
headers, so every call site reaches it either way.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* format: drop the trailing blank line left by the move
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <xroche@gmail.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Read sitemap files so URLs nothing links to are found
HTTrack finds URLs only by parsing links, so anything a site publishes solely
in its sitemap stayed invisible: robots.txt was already parsed, but its
Sitemap: lines were ignored and nothing else in the tree touched sitemaps.
Adds opt-in --sitemap (-%m), which probes the start host's robots.txt and
falls back to /sitemap.xml, and --sitemap-url (-%mu) for an explicit document.
Handles <urlset> and nested <sitemapindex>, plain or gzipped. Discovered URLs
enter with the full depth budget but still go through the wizard, so filters
and scope rules decide; a sitemap is not a filter bypass.
The parser reads attacker-controlled XML off the network, so it is capped on
URL count, index nesting, decompressed size and decompression ratio, and child
sitemaps must stay on the host that named them.
Closes#712
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* sitemap: distinguish the sitemapindex log line from a urlset one
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* sitemap: fix an out-of-bounds read, tighten the parser and the tests
lienrelatif() walked back from the last character of its current-path
argument without checking the path was non-empty, reading one byte before
the stack buffer. htsAddLink is the first caller to pass an empty savename,
which sitemap documents have because they are ingested rather than mirrored,
so ASan caught it on the new crawl test.
The parser drops a value whose numeric character reference decodes outside
printable ASCII, rather than leaving the reference verbatim and seeding a URL
the site never published, and classifies a document by its real root element,
so a comment naming the other one no longer flips urlset and sitemapindex.
The robots.txt line reader is bounded by the body size instead of relying on
a NUL terminator.
The self-test moves to 01_zlib-sitemap.test: MSan runs 01_engine-* only,
because an uninstrumented libz floods it with false positives.
Tests gain the assertions the earlier ones were missing: which of the
robots.txt route and the /sitemap.xml fallback was taken, that the sitemap
documents stay out of the mirror, that the off-host child sitemap is refused,
the sitemapindex nesting cap, the per-document URL cap at its production
value, and copy_htsopt coverage for the two new fields.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* sitemap: state what the decompression cap actually binds on
deflate tops out near 1032:1, so hts_codec_maxout never binds before the
64 MiB cap; the old comment implied a ratio guard that cannot fire.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* sitemap: a child sitemap is a fetch, so filters and robots.txt must gate it
An adversarial review found that a <sitemapindex> <loc>, and a robots.txt
Sitemap: line, went straight to hts_record_link: the request went out even
when a -* rule or robots.txt Disallow covered it. Only the <urlset> half ran
through the wizard. Gate the document itself on the filters and on
robots.txt, which is all that can apply: the wizard proper wants a referring
link, and its up/down travel rules would judge a child sitemap against the
parent sitemap's own directory. The robots.txt probe is exempt, being the
request that fetches the rules.
A 301 also used to end ingestion silently, since the engine re-queues the
target as a fresh link that carried no sitemap marking. That hit any site
redirecting http to https. The marking now follows the redirect.
The "N URL(s) added" counter reported what the scanner emitted rather than
what was taken, which hid both of the above; it now reads "N of M". The
fallback to /sitemap.xml keys on the same corrected count, so a robots.txt
whose only Sitemap: line is off-host or filtered still falls back. Root
classification skips a UTF-8 BOM and an XML namespace prefix, and the doc
list is cleared when a mirror starts rather than only when it ends.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* sitemap: gate each fetch by who asked for it, and anchor travel on the start URL
A nine-agent review found a scope escape: the sitemap document was its own
`premier`, so the wizard measured travel from wherever the site chose to put
its sitemap. A root /sitemap.xml therefore widened a /deep/dir/ crawl to the
whole host. The ingester now points the wizard at the crawl's own start link
and lets each seeded URL become its own anchor, which is what a command-line
seed gets.
Robots handling was both mistimed and undifferentiated. The Sitemap: lines are
now collected by robots_parse, on the same body in the same fetch, and acted on
after the parsed rules are installed rather than before; and the decision comes
from a new hts_robots_forbids extracted out of the wizard, so the sitemap path
inherits the -s1 filters-win override instead of a stricter hand-rolled check.
The four fetches are no longer treated alike: a sitemap the user names is user
intent, one the site declares invites the fetch, only the guessed /sitemap.xml
obeys a Disallow, and the URLs listed inside stay fully gated.
Also: hts_unescapeEntities replaces the private entity decoder, whose guard
tests and fuzz corpus it silently forfeited; hts_codec_head replaces hts_zhead,
which is only defined under HTS_USEZLIB; the composed URL buffer now fits two
maximal components plus a scheme, which a 2046-byte --sitemap-url reached; the
bounded search is promoted to htstools as hts_memstr; and the live state moves
from httrackp into htsoptstate, leaving two installed fields rather than three.
Tests gain the scope escape, the three robots cases, a cap-boundary control,
the handler invocation count and a compression-bomb decode. Every one was
checked against a deliberately broken build.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* sitemap: add the fuzz harness the parser was missing, and drop truncated Sitemap: lines
The file header called the scanner fuzzable while fuzz/ registered ten
harnesses and none for it. fuzz-sitemap feeds it raw XML, gzip-framed bodies
and truncated streams off a heap copy of exactly the input size, so an overread
is an ASan report rather than a quiet pass, with a four-file seed corpus.
60000 runs clean under ASan+UBSan.
robots_parse now drops a Sitemap: line that filled its scratch buffer instead
of handing on the half URL it was truncated to.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* fuzz: keep only the four sitemap seed inputs
A libFuzzer run writes its finds into the first corpus directory, and 191 of
them were committed with the harness.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* selftest: bound the sitemap document builders' snprintf accumulation
snprintf returns the length it wanted to write, so accumulating it blind
lets the next offset and size argument walk past the buffer. Guard each
step the way the argv builder above already does, and give the per-URL
loop a real remaining-space bound instead of a fixed 33.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* sitemap: date the new files 2026
The headers were copied from an existing file and kept its 1998 year.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* sitemap: keep the ingestion state out of htsoptstate
htsoptstate is embedded by value as httrackp.state, so a field at its tail
shifts every httrackp member declared after it: an offsetof probe put
warc_file at 141752 on master and 141760 on the branch. Move the pointer to
httrackp's own tail, where every existing offset holds and copy_htsopt still
ignores it.
Also renumber the crawl test to 89, master having taken 87 and 90, and give
the new option8 checkbox the hidden companion that 90_webhttrack-checkbox-clear
requires, plus its row in that test's table.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: renumber the sitemap crawl test to 95
Master took 88 through 93 and #720 claims 94.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* sitemap: realign the htsopt.h comments after the single-file merge
Master's longer LLint declarator moved the block's comment column.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* sitemap: translate the new GUI strings into the remaining 28 locales (#738)
The feature PR added the four LANG_SITEMAP* entries to lang.def with English
and Francais only; every other locale fell back to English in the WebHTTrack
form. Each file is written in its own declared charset.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
process_chain was incremented by the child in hts_entry_point(), so a caller
that spawned threads and immediately called htsthread_wait() saw a zero count
and returned at once, free to tear down state the children were about to read.
Count at spawn instead, under the same mutex the waiter reads.
httrack.c freed the option block before waiting; wait first.
Closes#747
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
--without-zlib only ever dropped -lz from LIBS; it never defined HTS_USEZLIB 0,
so every #if HTS_USEZLIB guard in the tree has been permanently true and the
build died with undefined references from minizip, htszlib.c, htswarc.c and
htsselftest.c. A zlib-free build is not reachable from there: the cache and the
WARC output are zip/gzip containers, and htsback.c already #error'd on
HTS_USEZLIB=0.
So make the requirement explicit. CHECK_ZLIB now errors out on --without-zlib
and on a missing header or library, keeping --with-zlib=DIR for a non-standard
prefix. htsback.c's #error moves to htsglobal.h where the knob is defined, with
a message that is accurate when it fires; its include of htszlib.h went with it,
unused.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Single-file output is MHT, which browsers no longer open
Add --single-file (-%Z): after the mirror completes, rewrite every saved
page in place with its stylesheets, scripts, images and fonts embedded as
data: URIs, while links between pages stay relative. The mirror remains a
browsable tree and each page also stands alone.
This cannot reuse the -%M path: MHT streams a MIME part per file as it is
saved, but a data: URI needs the asset's bytes when the page is written,
and pages are normally saved before their assets are fetched. The new pass
runs over the finished tree at the tail of httpmirror(), after the update
purge.
Audio, video, page-to-page links and anything over --single-file-max-size
(10 MB default) keep an ordinary link. References carrying a scheme are
skipped, which covers data: and makes a second --update run a no-op.
Resolution is clamped to the mirror root: the HTML is hostile input.
htsopt.h gains two tail-appended fields; VERSION_INFO is untouched.
Closes#713
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
* tests: renumber the single-file test to 87, run the self-test last
Master took 82 through 85 while this branch was open. The engine self-test
also moves to the end of the script: it and the crawl assertions cover
different ground, and failing first hid the crawl half.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
* Fix the parser and path bugs review found, and cover them
sf_relative_from() indexed one byte past from_dir's terminator whenever the
page's own directory was a prefix of the asset path, which is the ordinary
mirror layout: an out-of-bounds read that also miscounted the ../ prefix and
emitted a broken link. Two review agents hit the same ASan trace; the branch
had no test because every nested asset in the fixture was under the cap.
Also from review: an escaped quote no longer ends a CSS string early (which
exposed its contents to the url()/@import scanner), @import url(...) now
inlines like the quoted form, a raw-text element ends only on a real end tag
rather than any prefix of one, an over-wide tag is copied through by the same
quote-aware scan instead of a second quote-blind one, <!--> is an empty
comment, whitespace before a tag's > survives, and a page cannot inline more
than SINGLEFILE_MAX_PAGE_SIZE, which bounds the multiplicative @import
fan-out. A failed encode no longer leaves a payload-less data: prefix behind,
and base_dir is matched against the root on a component boundary.
sf_readfile now wraps a new readfile2_utf8() rather than being a fifth copy
of the readfile family. Docs place --single-file beside -%M instead of
implying it supersedes it: MHT stays the better container, this wins on
opening anywhere.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
* singlefile: tighten two comments
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
* tests: cover the option plumbing review found untested
copy_htsopt's two new fields, including the >0 guard that must not let an
unset source clear the target's default; -%Z0; and a rejected
--single-file-max-size argument falling back to the built-in cap.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
* tests: stand the single-file GUI check down where htsserver is absent
The Windows job builds httrack.exe only, and reaches this file through its
*_local-*.test glob, so requiring htsserver failed the whole test there even
though every crawl assertion had passed. Skip that half instead, and run the
engine self-test before it so the skip cannot swallow it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
* tests: check the sprintfbuff master just made warn_unused_result
#722 gave slprintfbuff the attribute, so the over-wide-tag fixture's call
became the one warning this branch adds over master's baseline. Handle it the
way that PR's own self-test code does; the buffer cannot truncate here.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
* single-file: translate the new GUI strings into the remaining 28 locales (#739)
The feature PR added the four LANG_SINGLEFILE* entries to lang.def with English
and Francais only; every other locale fell back to English in the WebHTTrack
form. Each file is written in its own declared charset.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* single-file: restore the mirror file mode, and give the guards real coverage
The rewrite spools to a .sfnew temp and renames it over the page, which
bypasses filecreate() and with it the HTS_ACCESS_FILE chmod the engine puts on
every other mirrored file. Under a restrictive umask the pages came out 0600
while their assets stayed 0644, so a mirror served by a webserver or shared
with a group lost read access on exactly the pages. chmod the spool before the
rename.
Two guards had no coverage: deleting the scheme/data: check in sf_resolve left
both the self-test and the crawl test passing, because the fixtures resolved to
paths that were absent either way, and the per-page inline budget was never
exercised. The fixtures now plant a file where each guard's removal would land
the walk, and a self-importing stylesheet measures the budget against a
large-budget control. Charging that budget after the nested rewrite instead of
before let an @import chain spend what its ancestors had already claimed and
drove it negative; charge it up front and refund on failure.
The attribute table missed the lazy-loading attributes hts_detect[] already
downloads, so a modern page inlined almost nothing: add data-src, data-srcset,
lowsrc, object@data and embed@src, and record why the rest stay links.
The new web GUI checkbox had no hidden companion input, so it could be ticked
but never cleared, which is what #725 fixed for every other box. Master's test
90 catches it once the branch merges.
Adds a libFuzzer harness over the rewriter, since it re-serializes hostile
HTML, and renames the crawl test to 91 now that master holds 87 and 90.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* single-file: document the inlined-stylesheet limitation in the CLI guide
Recorded in htssinglefile.h already; the user-facing guide is where someone
raising --single-file-max-size will look.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: move the Windows note down to the gate it explains
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* ci: register the single-file test's Windows skip
Its GUI half needs htsserver, which the Windows job does not build, so the test
now exits 77 there instead of reporting a pass for assertions it never ran. The
skip list is pinned, so it has to be declared. Both Windows jobs reported
fail=0; only the list check was red.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <xroche@gmail.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Report what a crawl changed against the previous mirror (--changes)
--update already knows which resources were new, which changed and which the
server called unchanged, and throws it away: the flags reach file_notify() and
go no further than a log line, while deletions exist only as a side effect of
purging. --changes (-%d) keeps all of it and writes hts-changes.json plus a
one-line summary in the log.
"Changed" means the bytes differ, not that the server re-sent the resource.
Comparing the mirrored files directly would not work: HTTrack stamps every
parsed page with the crawl date via the footer, so those bytes differ on every
run. Payloads are compared instead, the previous one coming from the cache for
parsed pages and from the local copy sampled just before it is overwritten for
everything else.
The mirror-relative path, not the URL, is the accumulator's key, so a redirect
and its target that share a save name are one entry; and only the first notify
for a file samples its pre-run state, so a retried transfer is not counted
twice. What counts as already mirrored comes from the previous run's file
index rather than from the file's presence on disk: a partial left by this
crawl's own failed attempt is on disk but was never part of the previous
mirror.
The deleted set is now computed whether or not purging is enabled; unlinking
still happens only under --purge-old.
Closes#714
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
* changes: never leave a stale report behind
A crawl that mirrored nothing created no accumulator, so hts_changes_close_opt
returned without writing and the previous run's report stayed on disk as if it
described this one. Write it whenever --changes is on. The no-data rollback is
the deliberate exception, and is now documented: it restores the previous cache
generation, so leaving the matching report alone is the consistent behaviour.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
* changes: drop em dashes from the format page
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
* changes: fix the review's three findings
The size shortcut compared rendered on-disk sizes even for parsed pages, whose
payload digests describe something else entirely, so it decided the outcome
before the payload comparison could run: a page whose payload never changed but
whose rewritten links moved read as changed. It now only applies when both
digests describe the file on disk.
file_notify() reaches the accumulator from the FTP download thread as well as
the main one, and the lazy allocation, the coucal write and the entries realloc
were all unguarded. Every entry point now takes a mutex, and the HTML hook does
its cache read before taking it, since that read can itself re-enter
file_notify() and move the array.
Two fixtures cover what nothing did: a gzipped direct-to-disk body that changes
at constant length (without the pre-sample before the decoded temp is renamed,
it reads as unchanged), and a page with a fixed payload behind a redirect whose
target is renamed, so its file on disk changes length while its bytes do not.
Both were checked against builds with the respective fix reverted.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
* changes: document the renamed-file case in the format page
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
* changes: refresh two stale test comments
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
* Merge origin/master into feat/change-report
Both sides appended to tests/Makefile.am's TESTS; kept master's
86_local-proxytrack-cache-longfields.test alongside 88_local-changes.test.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
* changes: translate the new GUI strings into the remaining 28 locales (#740)
The feature PR added the two LANG_CHANGES* entries to lang.def with English and
Francais only; every other locale fell back to English in the WebHTTrack form.
Each file is written in its own declared charset.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* changes: fix the review's blocking findings
Lock the report path against the FTP thread the crawl never joins, seal the
accumulator once the report is written, key entries off the project directory
so the report survives --cache=0, stop calling a file gone when the crawl only
failed to re-fetch it, and skip the hook's work entirely when --changes is off.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* changes: prove the fixes, and give the report a cache-off mode
Adds a changes-race self-test (the FTP shape: notifier threads against the
report path), fixtures for a transfer the crawl never completes, for a leftover
file at a name the crawl mirrors fresh, and for a cache-off mirror, plus a pass
with purging on. Registers the web GUI's --changes box with the clearing
companion master's #725 now requires, and documents the degraded mode.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* changes: keep the failed-transfer case portable
A connection killed before the status line surfaces differently on macOS, where
it truncates the mirrored file to zero (#748). Assert only what holds on both:
it is never reported gone, and its file survives.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <xroche@gmail.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both binput calls pass HTS_URLMAXSIZE, but they write into one
line[HTS_URLMAXSIZE * 2] and binput puts its NUL at s[max]. A first field
that fills its whole bound leaves the second writing line[2048], one past
the array. ASan reports a stack-buffer-overflow on a crafted .ndx. Bound
the second by what the first left.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* PT_GetTime's failure fallback formats as day 00
An all-zero struct tm has tm_mday == 0, and the ARC filedesc line prints
tm_mday raw, so a gmtime failure would emit "...0100" where the day
belongs. Use the epoch, as PT_SaveCache__Arc_Fun already does for an
unparseable Last-Modified.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* The fallback is reachable, so test it
I claimed no reachable trigger. Wrong: file_timestamp() passes st_mtime
through untouched, and a cache whose mtime is past gmtime's range takes
the fallback. Only the ARC loader is safe, because it overwrites the
timestamp with the filedesc line's own 4-digit-year date.
Test 88 sets an out-of-range mtime on a zip cache and pins the emitted
date; without the fix it reads 19000100000000.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Register the Windows skip, and make the skip path real
The Windows job pins an exact expected-skip list, so a new conditional
skip fails it; test 88 skips there because NTFS will not hold an mtime
past gmtime's range. Its own skip logic was also dead code: "rc=0 || rc=$?"
never runs the right-hand side, and only set -e was carrying python's 77
through. Capture the status properly and tell a clamp apart from a real
failure.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Fourteen sites, all safe today: literals into sized fields, and two
same-sized array copies. The three writing "" through r->location, a
char *, become a direct terminator, since strcpybuff would have taken its
pointer path and lost the bound.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Fold the three copies of the clip idiom into strclipbuff()
htscache.c and the two readers in proxy/store.c each spelled out
clear-then-strlncatbuff, in binaries that share no code. A helper in
htssafe.h states the contract once, and evaluates its arguments once: the
macro form expanded refvalue and refvalue_size twice, and (refvalue_size)
- 1 would have wrapped to SIZE_MAX had any call site ever passed 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Pin the capacity-2 boundary
The cases jumped from the degenerate capacity 1 straight to 8, so a
defect confined to small-but-not-degenerate sizes passed: clipping a
two-byte destination to the empty string instead of one character.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* htsserver: a request naming a directory spins the server forever
fopen() succeeds on a directory on POSIX and every read from it fails with
EISDIR without ever raising EOF, so smallserver()'s "while (!feof(fp))" serving
loop never terminates. GET /server/ needs no session id and no project to reach
it, and the accept loop is single-threaded, so one unauthenticated local request
wedges WebHTTrack for good.
Refuse a directory before fopen() so the 404 branch answers, rather than only
ending the loop: a loop-only fix would serve every directory as an empty 200.
The two loops fed a client-influenced path also stop on ferror(), since a read
that fails for any other reason spins the same way.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* htsserver: whitelist regular files instead of blacklisting directories
fexist() (htsserver.h) is the stat + S_ISREG predicate this file already uses
for the "file-exists:" template op, so gate the serving fopen() on it rather
than on a fresh negated is_directory(). Refusing only directories still handed
FIFOs to the same code path, where fopen() blocks with no writer and kills the
single-threaded accept loop for good.
Test 91 gains the FIFO case and a POST that loads a project whose
hts-cache/winprofile.ini is a directory: that fopen() sits ahead of every guard,
so it covers the ferror() check on the project-load read loop.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* htsserver: fold the two comments above the serve guard into one
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* htsserver: clip the "save settings" failure messages
The three sprintf(tmp[1024], ...) sites reporting a failed profile save
quote a project path composed from the posted "path" and "projname"
fields, neither of them bounded. A sid-authenticated save with a
1200-byte path smashes the stack buffer and takes the server down.
Route them through a SET_ERRORF() that formats into a fixed buffer and
absorbs the truncation once, so the message clips instead of aborting:
the text comes from the client, and a quieter denial of service is no
better than a loud one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: reach the second message without a 1030-byte mkdir
The tree the case needed came within a few bytes of macOS's PATH_MAX
once /var resolves to /private/var. A symlinked hts-cache gets there
instead: structcheck() lets a non-directory through, and opening
winprofile.ini under it fails with ENOTDIR whatever the uid.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* htssafe: one clipping printf helper for the three failf wrappers
htsblk_failf(), PT_Element_failf() and htsserver's format_error() all had
the same body; slprintfbuff_clip() now owns it. A (void) cast on
slprintfbuff() is no substitute: GCC warns through warn_unused_result.
vslprintfbuff() also empties dest before formatting, so a vsnprintf that
fails outright cannot leave the caller publishing uninitialized stack.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: assert the clipped length, not just the message text
Two of the three sites only checked that the message appeared, which the
old sprintf did equally well; each now compares the rendered message
against the 1023-byte clip. The failing-write branch no longer needs
/dev/full either: the server runs under a file-size limit, so macOS gets
the same coverage.
Renumbered 86 to 89, which no other pending change claims.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: build the oversized profile without a 16k printf width
The macOS leg reached the third message with no error set at all, so the
write the file-size limit was supposed to break had gone through. Build the
profile by doubling a 1024-wide printf, assert its length, and assert the
init file came out short, so a limit that does not bite names itself
instead of surfacing as a missing message.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: build the long paths from the physical temp directory
That is what broke the macOS leg: TMPDIR there resolves through the /var
symlink, so the 1020-byte init file the third case opens was really 1028
bytes to the kernel and the open failed with ENAMETOOLONG. Both cases sit
within a few bytes of macOS's 1024-byte PATH_MAX, so the paths have to be
measured after resolution.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Options ticked on by default cannot be un-ticked in the web GUI
An unchecked HTML checkbox posts nothing, so htsserver never overwrites the
value it already holds. Every box in the wizard is one-way: once the stored
value is "1", whether htsserver seeded it at startup, a loaded profile set it,
or the user ticked it earlier in the session, un-ticking and submitting leaves
the option on and draws the box ticked again. Only four boxes, all in
option1.html, carried the companion hidden field that guards against this.
Add it to every remaining bare checkbox, and switch cookies and parsejava to
${ztest:...} so a cleared box emits --cookies=0 / --parse-java=0; ${test:...}
renders nothing at all when the value is empty, which is not "off" for an
option the engine turns on by default.
index, urlhack and keep-alive are deliberately left alone: their long options
are declared "single" in htsalias.c and optalias_check drops the =value, so
--index=0 resolves to -I and turns the option back on. That parser bug is
pre-existing and needs its own fix.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: pin last-write-wins and cover every checkbox
The runtime leg posted each name once, so a first-wins body parser would
have passed while the fix did nothing in a real browser: post the
duplicated name in both orders and assert the last value wins. Replace the
four hand-written option cases with a table covering all 28 non-skipped
boxes, asserting the command-line token and the Windows-profile key each
state emits, plus a completeness check so a new box cannot slip through
unexercised.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: rename the loop variable shadowing the scanned page
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* tests: a failed request must not read as a clean security verdict
Under pipefail, "request | grep -q MARKER && fail" skips the fail when the
request itself errors: the leak checks in tests 78 and 85 then pass without
ever having run. Capture the reply first and fail loudly if it never arrived.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* AGENTS.md: record the fail-open assertion shape
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: a reply that proves nothing must not read as a clean verdict
The previous commit converted some of the fail-open assertions and left three.
78's refusal loop still piped into "grep -q ... && fail": grep -q exits on the
first match and SIGPIPEs the producer, so under pipefail a hostile reply that
pads its Location past the 64 KB pipe buffer suppresses the failure exactly as
a dead probe would. 85's fetch() only required a non-empty reply, so a
truncated body or a 302 to the file passed the leak checks marker-free, and no
assertion looked at the status line at all. 78's store probe had no emptiness
guard, so an empty page read as "the store was not written".
Match from here-strings throughout, give fetch() the status each caller
expects, and route 78's store probe through a helper that requires a served
page. 77's X-Injected check had the same shape.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The trim that walks back to the last '/' starts at `curr + strlen(curr) - 1`,
which is `curr - 1` when the path is empty. The loop then dereferences it.
An empty path is reachable today: the pre-pass that strips a query does
`strncatbuff(newcurr_fil, curr_fil, a - curr_fil)`, so any `curr_fil` starting
with '?' hands the walk an empty string. `-#test=relative "dir/page.html" "?x"`
under ASan reports the underflow.
The read is one byte and the loop stops immediately either way, so the guard
changes no output: over the 484 ordered pairs of a 22-value path corpus, run
against builds that force the byte before the buffer to 0 and to '/', the
guarded and unguarded results are identical.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Five PRs across the engine and ProxyTrack turned up the same few traps
more than once, and none of them are obvious from the code.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Clear the last three compiler warnings
finalurl was sized for one of the two URLs it concatenates. The IIS-bug
example callback overwrites a suffix in place with a same-length
replacement and must not terminate the string, which is memcpy, not
strncpy. The coucal bench's if/else chain has no final else, so result was
only initialised on the paths gcc could not prove exhaustive.
A clean build now reports zero warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Bound the IIS suffix copy by what matched, not by the table
Copying strlen(replacement) leaves the "MUST be the same sizes" comment as
the only thing standing between a future table edit and an overflow. j is
the number of bytes just matched in the destination, so copying j is safe
whatever the table holds.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* A cache field wider than ours aborts the engine instead of clipping
The read-side ZIP_READFIELD_STRING used strlcpybuff, and the whole *_safe_
family aborts on overflow rather than truncating. Since the header line is
bounded only by HTS_URLMAXSIZE and msg is 80 bytes, a cache written by
another build, or a corrupt one, kills the crawl outright. The corrupt-cache
self-test already promises "rejected per-entry, never crash".
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Pin the clip to each field's own capacity
Review found one case exercised only msg[80], so a hardcoded clip length
passed. lastmodified[64] is narrower, and no single constant satisfies
both. The forged replacement was also one byte longer than the line it
overwrote, which only worked because corrupt_patch copies exactly the
pattern length.
Also stop claiming another build's cache can trigger this: the writer
emits each field from the same struct the reader fills, so it takes a
corrupt cache.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
t_StatsBuffer was defined twice, byte for byte, in httrack.h and htsweb.h,
with NStatsBuffer duplicated alongside. httrack and htsserver each keep
their own array, so nothing catches the two drifting apart, and the last
change to the struct had to be applied to both by hand. Both now include
src/htsstats.h; sizes and offsets are unchanged.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Cached entry with no usable date crashes the ARC writer
PT_SaveCache__Arc_Fun dereferenced convert_time_rfc822() straight into the
record line, so any entry whose Last-Modified is absent or unparseable took
proxytrack --convert down. The sibling caller a thousand lines up already
guards the same call; this one fills in the epoch instead.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Assert the archive date, not just the entry
Review found the test blind to the two mutants that matter: a guard firing
unconditionally clobbers every valid date to the epoch, and one that skips
the year and day emits a month and day of 00. Grepping only for the URL saw
neither. Assert the date field exactly, and add a valid-date case so the
untouched path is pinned too.
A bare "Last-Modified: 0" crashes the same way, so it joins the cases.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Bound the remaining writes into htsblk.msg[80]
Nine writers still filled the 80-byte msg with no bound. Six sprintf the
result of strerror(), whose longest glibc string is 49 bytes in the C
locale and 72 in fr_FR against 46 bytes of room after the longest prefix.
Realistic connect() errnos still fit, so this was latent rather than live.
The FTP helper's .ok result file had no excuse: it was copied byte by byte
until EOF into the same 80 bytes. Split that parse out as
back_read_ftp_result() so a self-test can drive it, and stop at capacity.
These sites never appeared in the -Wformat-truncation cluster that added
htsblk_failf: the diagnostic only fires on a bounded snprintf whose return
is discarded, so a raw sprintf is invisible to it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Close the self-test's blind spots around msg[]
Six of eight mutants survived the first version. The neighbour canary
compared against zero, so it saw a stray 'X' but not the stray NUL an
off-by-one terminator actually writes; poison it instead. Only the
over-capacity case existed, so padding every message to 79 bytes or eating
its last character both passed, as did a sign-extended 0xff reading as EOF
and the new unparseable-status branch, which had no coverage at all.
Same zero-comparison weakness applied to the htsblk_failf canary inherited
from #715, so that one is poisoned too.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Bound ProxyTrack's cache header copies
ZIP_READFIELD_STRING in proxy/store.c took no destination size and used a
raw strcpy, while the engine's namesake in htscache.c has taken a
refvalue_size for years. The source is a cache-entry header line bounded
only by line[HTS_URLMAXSIZE + 2].
contenttype[64] sits just before the location pointer in the calloc'd
_PT_Element, so an over-long Content-Type walks over charset and then
over location, which the next Location: line copies through.
Clip rather than reject: the engine stores Content-Type in 128 bytes, so
a valid cache legitimately carries fields wider than ours.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Bound the ARC reader's header copies too
Review found store.c carries a second copy of the same macro:
HTTP_READFIELD_STRING feeds the ARC reader from index->line[2048], twice
the ZIP path's reach, into the same contenttype[64] and its neighbouring
location pointer. proxytrack --convert on a plain-text ARC segfaults.
Test 86 covered one of eight fields, so a per-field clip and a
one-size-fits-all clip were indistinguishable. It now overshoots every
destination and asserts each surviving length against its own capacity,
and exercises the ARC reader without needing python.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Fold htsblk/PT_Element failure messages into a clipping helper
The 18 remaining -Wformat-truncation warnings all came from building a
diagnostic string out of a remote server's reply and dropping snprintf's
return. Add htsblk_failf() for htsblk.msg[80] and PT_Element_failf() for
ProxyTrack's msg[1024]: both clip to fit and absorb the discard once, so
the obligation is not silently laundered at each call site.
msg[80] lives in the installed htsopt.h, so growing it would break the
ABI; a clipped FTP banner is the intended outcome anyway. The display
StatsBuffer.state is not installed, so it grows to fit back->info instead.
Also bounds two unbounded sprintf(r->msg, ...) in ProxyTrack's cache
reader and moves the raw strcpy neighbours to strcpybuff.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Catch a one-past-the-end write into msg[]'s neighbour
Review found the new self-test blind to a store at msg[sizeof(msg)]: it
lands in the adjacent contenttype field, which no assertion read, and an
intra-struct overflow is invisible to ASan and _FORTIFY_SOURCE. Check the
neighbour after every call, and add the exact-fit case the block lacked.
Correct the contract comment too: msg is not purely diagnostic, it
round-trips through the cache as X-StatusMessage.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* tests: arm test 82's mirror root through a profile save
#707 made /website/ serve only the root htsserver recorded at structcheck
success, so the posted projpath test 82 relied on no longer names anything.
Each PR was green alone; the pair was not.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: drop the now-unused project-path argument
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* htsserver builds the redirect Location header in a 256-byte stack buffer
The POST redirect path checks strlen(file) but sprintf's newfile, which comes
straight from the client's "redirect" POST field with no length cap. A 300-byte
value overflows tmp[256]. The same value reached the Location header with no
CR/LF check, so it could split the response and inject headers.
Append into the dynamic String the other headers already use, and drop the
header entirely when the value carries a CR or LF.
The listen socket was SOCaddr_initany, so the server answered the LAN and not
just the local browser it exists to serve. Bind 127.0.0.1 by default, with
--bind <addr> to widen it again, resolved through the existing gethost() helper
the way proxytrack already does.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: satisfy shellcheck and shfmt in the new server test
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: drop the pre-fix narration from the oversized-value comment
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: assert the bound socket, not the announced URL
The listen-address assertions only compared the URL= banner, which is a
literal echo of argv: a build that announced 127.0.0.1 while binding the
wildcard passed. Probe 127.0.0.2 on the same port instead, which a wildcard
listener takes and a loopback-only one leaves free.
Also refuse an empty --bind, which fell through to every interface and
silently undid the new default.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: bound the response read and the empty --bind run
The recv() loop had no timeout and read until EOF; htsserver need not close
the connection after responding, which wedged the macOS runner for over an
hour. Stop at the end of the header block, which is all the test reads.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* htsserver: the session id must gate the request body, not the reply
Every field of a POST body is written straight into the one global key store
the templates and the command dispatcher both read, and "command" from there
reaches the engine. The gate ran after that write and compared "sid" against
"_sid" -- but "_sid" is copied into "sid" beforehand so the templates can
render it, so a request that simply omitted the field compared equal to
itself. Only a wrong id was refused; an absent one passed. Clearing the reply
afterwards does not help either, because the dispatcher sits outside the reply
guard.
Authenticate before parsing instead: scan the raw body for "sid", require at
least one occurrence and reject if any of them differs, and drop the body
untouched when it does not match. That leaves the shared template key alone,
and it closes the dispatcher for free.
A refused request also emitted only a Content-length line, since the status
line for that branch was behind _DEBUG. Any client reads that as a protocol
error, which is how test 68 failed rather than reporting the refusal. Send a
403 instead.
Tests 68 and 77 posted without an id, which is what the engine used to accept,
so both now fetch the one the server renders into the form. Test 78 covers
accept, missing, empty and wrong, asserts the 403, and probes the key store
through ${projname} rather than the suppressed reply -- a reply-only assertion
passes even when the write goes through.
Also fix a leak that hung macOS CI: start() runs inside a command
substitution, so its $! never reached the parent and stop() guarded on an
empty variable, leaving one htsserver per call. Test 77 starts four, which is
exactly the four orphans the runner reported while sitting for half an hour
behind a green test log. Take the pid from the PID= line the server already
announces.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* htsserver: serve the crawled mirror verbatim, never through the expander
WebHTTrack serves the mirror under /website/ from the same small server as its
own GUI, and the decision to run a response through the ${...} template
expander was a substring test for ".htm" on the request path. A mirrored page
therefore had its directives evaluated: ${_sid} rendered the live session id,
handing the crawled site the token that authenticates commands on the local
GUI, and ${do:...} gave it the rest of the template verbs.
The /website/ prefix was already detected, but only to keep the crawl-state
override from hijacking a mirror request. Reuse it as the expansion gate, so
expansion is limited to files under the GUI's html root, and re-evaluate it
after that override, which can substitute a GUI page for a mirror path.
Mirrored pages keep their text/html type: verbatim must not turn browsing the
mirror into a download.
Note that mirrored content still shares the control origin, so a script in it
can read the session id from a GUI page itself; that is a separate fix.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* htsserver: stop letting a posted projpath name the /website/ root
The static-file path in smallserver() was composed with a length guard that
measured only the server root and the request path, then wrote the posted
"projpath" field into a 1024-byte stack buffer: a 2000-byte projpath followed
by any /website/ request smashed the stack (buffer overflow detected, server
gone). The guard also summed two untrusted lengths before comparing, the shape
that can wrap and pass.
Composition now goes through a bounded, non-aborting append that keeps the
untrusted length alone on one side, and /website/ is served from the project
directory the server itself set up, rejecting a ".." in it, rather than from
whatever root the request body claimed. Without that, projpath=/etc/ plus
GET /website/passwd read an arbitrary file, since the ".." check looked at the
request path only. fsfile is also cleared before the error-redirect branch,
which could otherwise reach fopen() uninitialized.
tests/80_webhttrack-projpath.test drives the three cases against a live server
and keeps a legitimate project browsable as the control.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: establish test 79's mirror root through the server
/website/ no longer serves the posted projpath, so injecting one as a
fixture stopped working and test 79 got a 404 instead of the mirrored
page. Save a profile first (no command_do=start, so nothing crawls) to
make the server record the root, then plant the file under it.
Re-checked against a reverted expander gate: still fails there, so the
fixture change did not cost the assertion its teeth.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: cover the running-crawl half of the /website/ override
Test 84 only exercised the idle server, where the override never fires and
the recomputed virtualpath is indistinguishable from the stale one. Drive a
crawl through the server so /website/*.html is rewritten to the GUI refresh
page, which 404s without the recompute.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: cover the overflow, the '..' rejection and the fsfile hoist
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: fix the merged test 84 duplicate post() helper
The rename-side and the branch-side each defined post(), with different
argument shapes; the last one won and silently mangled the save body.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: trim the new comments
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: drop the stale start() arg comments in test 84
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* WebHTTrack's mirror links are dead file:// URLs
The GUI is served from http://127.0.0.1:PORT/, and browsers refuse to navigate
from an http: page to a file: URL, so the "browse the mirror" links in
finished.html and file.html did nothing when clicked. They now point at the
server's own /website/ route, which already serves the project directory. The
desktop entry's browse mode still works, because there the shell hands the
file: URL to the browser instead of navigating from a page.
The per-project picker in file.html goes too: /website/ is bound to the running
project, and reaching the other projects over HTTP would mean serving the whole
mirror tree from the control origin.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* finished.html: link the served mirror, not a dead file:// URL
The page is served over http:, so its file:// link is a cross-scheme
navigation that Chrome and Firefox refuse. Nothing happened on click. The
mirror is already reachable at /website/, which the two list entries just
below were using all along.
file.html keeps its file:// links. It is the fresh-session entry point for
sites mirrored earlier, and /website/ is bound to one project, so pointing
it there would trade a dead link for a 404. Serving an arbitrary past
project needs a route that does not exist yet.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: tighten the browse-link assertion and trim its comments
The runtime check for href="/website/index.html" also matched the list
entry below the anchor, so it passed on the unfixed page; match the
anchor by its mirror-path label instead.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* htsserver builds the redirect Location header in a 256-byte stack buffer
The POST redirect path checks strlen(file) but sprintf's newfile, which comes
straight from the client's "redirect" POST field with no length cap. A 300-byte
value overflows tmp[256]. The same value reached the Location header with no
CR/LF check, so it could split the response and inject headers.
Append into the dynamic String the other headers already use, and drop the
header entirely when the value carries a CR or LF.
The listen socket was SOCaddr_initany, so the server answered the LAN and not
just the local browser it exists to serve. Bind 127.0.0.1 by default, with
--bind <addr> to widen it again, resolved through the existing gethost() helper
the way proxytrack already does.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: satisfy shellcheck and shfmt in the new server test
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: drop the pre-fix narration from the oversized-value comment
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: assert the bound socket, not the announced URL
The listen-address assertions only compared the URL= banner, which is a
literal echo of argv: a build that announced 127.0.0.1 while binding the
wildcard passed. Probe 127.0.0.2 on the same port instead, which a wildcard
listener takes and a loopback-only one leaves free.
Also refuse an empty --bind, which fell through to every interface and
silently undid the new default.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: bound the response read and the empty --bind run
The recv() loop had no timeout and read until EOF; htsserver need not close
the connection after responding, which wedged the macOS runner for over an
hour. Stop at the end of the header block, which is all the test reads.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* htsserver: the session id must gate the request body, not the reply
Every field of a POST body is written straight into the one global key store
the templates and the command dispatcher both read, and "command" from there
reaches the engine. The gate ran after that write and compared "sid" against
"_sid" -- but "_sid" is copied into "sid" beforehand so the templates can
render it, so a request that simply omitted the field compared equal to
itself. Only a wrong id was refused; an absent one passed. Clearing the reply
afterwards does not help either, because the dispatcher sits outside the reply
guard.
Authenticate before parsing instead: scan the raw body for "sid", require at
least one occurrence and reject if any of them differs, and drop the body
untouched when it does not match. That leaves the shared template key alone,
and it closes the dispatcher for free.
A refused request also emitted only a Content-length line, since the status
line for that branch was behind _DEBUG. Any client reads that as a protocol
error, which is how test 68 failed rather than reporting the refusal. Send a
403 instead.
Tests 68 and 77 posted without an id, which is what the engine used to accept,
so both now fetch the one the server renders into the form. Test 78 covers
accept, missing, empty and wrong, asserts the 403, and probes the key store
through ${projname} rather than the suppressed reply -- a reply-only assertion
passes even when the write goes through.
Also fix a leak that hung macOS CI: start() runs inside a command
substitution, so its $! never reached the parent and stop() guarded on an
empty variable, leaving one htsserver per call. Test 77 starts four, which is
exactly the four orphans the runner reported while sitting for half an hour
behind a green test log. Take the pid from the PID= line the server already
announces.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* htsserver: serve the crawled mirror verbatim, never through the expander
WebHTTrack serves the mirror under /website/ from the same small server as its
own GUI, and the decision to run a response through the ${...} template
expander was a substring test for ".htm" on the request path. A mirrored page
therefore had its directives evaluated: ${_sid} rendered the live session id,
handing the crawled site the token that authenticates commands on the local
GUI, and ${do:...} gave it the rest of the template verbs.
The /website/ prefix was already detected, but only to keep the crawl-state
override from hijacking a mirror request. Reuse it as the expansion gate, so
expansion is limited to files under the GUI's html root, and re-evaluate it
after that override, which can substitute a GUI page for a mirror path.
Mirrored pages keep their text/html type: verbatim must not turn browsing the
mirror into a download.
Note that mirrored content still shares the control origin, so a script in it
can read the session id from a GUI page itself; that is a separate fix.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: cover the running-crawl half of the /website/ override
Test 84 only exercised the idle server, where the override never fires and
the recomputed virtualpath is indistinguishable from the stale one. Drive a
crawl through the server so /website/*.html is rewritten to the GUI refresh
page, which 404s without the recompute.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* webhttrack: "max site size" set a per-file cap instead of the overall one
step4.html mapped all three size fields of the wizard onto --max-files (-m),
so "Max site size" emitted a per-file limit rather than --max-size (-M), and
landed a second -m on the command line. That second -m also clobbered the HTML
per-file limit: a bare -m<n> resets maxfile_html, so whichever of the two came
last won. Point sizemax at --max-size and emit the bare -m before the -m,<n>
form so both per-file caps survive.
Two template typos in the same family, where a malformed ${...} renders as
nothing or as its own key instead of erroring: the winprofile.ini writer's
Dos=${dos was missing its closing brace, and option2b.html's OK button read
${LANG_OK] with a bracket.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* tests: carry the session id in test 81's POST
The gate that landed with #700 refuses a body without one, so the wizard
POST came back refused and the option audit had nothing to read.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: close the confirmation-biased gaps in test 81
Post the size fields empty too, so a step4.html that lost its ${test:} guard
and rendered a valueless --max-files= is caught; assert the winprofile.ini
MaxHtml/MaxOther/MaxAll keys the header claimed to audit; and pin the OK button
label, since an unknown ${LANG_} key renders empty and passed the absence check.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds `sprintfbuff()`/`slprintfbuff()` to `htssafe.h`: a formatted print that truncates to fit and returns whether it had to, marked `warn_unused_result` so the answer cannot be dropped. It fills the gap between the `strcpybuff` family, which aborts on overflow, and `String`, which grows without bound. Abort is the wrong contract wherever the text is built from a remote peer's reply.
Four `-Wformat-truncation=` sites used the result as if `snprintf` had never truncated. Two only needed a bigger destination, so they get one: `hts_finish_makeindex`'s `tempo` was a flat 1024 against a 2048-byte escaped URL and is now sized off it, and the wizard's `cmd[4096]` could not hold the answers it concatenates. The other two cannot grow. `create_back_tmpfile` formats `<url_sav>.bak` into a buffer the same size as `url_sav`, and that struct is installed, so a dropped extension would alias the backup onto the live file that `back_finalize_backup()` unlinks. ProxyTrack's `startUrl[1024]` is fed by cache content. Both take the error path they already had, and ProxyTrack moves to the next cache entry rather than publishing a clipped one.
On what the wrapper buys, since it is not what I first assumed: checking a raw `snprintf` return inline silences the warning just as well. The wrapper's value is that the capacity comes from `sizeof`, the check is the default rather than the exception, and `warn_unused_result` makes skipping it visible.
`-#test=strsafe` covers the primitive (exact fit, one over, 4 KB source, `size == 1`, trailing canary, destination repoisoned between cases), and `-#test=makeindex` gains a first link whose escaped form overruns the old buffer. Both were mutation-checked. The ProxyTrack change ships without a direct test: its only observable is the catalog page, which renders solely as a PROPFIND fallback I could not drive from curl. 25 gcc warnings down to 21; the rest of the cluster is diagnostic-only and follows separately.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* htsserver builds the redirect Location header in a 256-byte stack buffer
The POST redirect path checks strlen(file) but sprintf's newfile, which comes
straight from the client's "redirect" POST field with no length cap. A 300-byte
value overflows tmp[256]. The same value reached the Location header with no
CR/LF check, so it could split the response and inject headers.
Append into the dynamic String the other headers already use, and drop the
header entirely when the value carries a CR or LF.
The listen socket was SOCaddr_initany, so the server answered the LAN and not
just the local browser it exists to serve. Bind 127.0.0.1 by default, with
--bind <addr> to widen it again, resolved through the existing gethost() helper
the way proxytrack already does.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: satisfy shellcheck and shfmt in the new server test
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: drop the pre-fix narration from the oversized-value comment
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: assert the bound socket, not the announced URL
The listen-address assertions only compared the URL= banner, which is a
literal echo of argv: a build that announced 127.0.0.1 while binding the
wildcard passed. Probe 127.0.0.2 on the same port instead, which a wildcard
listener takes and a loopback-only one leaves free.
Also refuse an empty --bind, which fell through to every interface and
silently undid the new default.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: bound the response read and the empty --bind run
The recv() loop had no timeout and read until EOF; htsserver need not close
the connection after responding, which wedged the macOS runner for over an
hour. Stop at the end of the header block, which is all the test reads.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* htsserver: the session id must gate the request body, not the reply
Every field of a POST body is written straight into the one global key store
the templates and the command dispatcher both read, and "command" from there
reaches the engine. The gate ran after that write and compared "sid" against
"_sid" -- but "_sid" is copied into "sid" beforehand so the templates can
render it, so a request that simply omitted the field compared equal to
itself. Only a wrong id was refused; an absent one passed. Clearing the reply
afterwards does not help either, because the dispatcher sits outside the reply
guard.
Authenticate before parsing instead: scan the raw body for "sid", require at
least one occurrence and reject if any of them differs, and drop the body
untouched when it does not match. That leaves the shared template key alone,
and it closes the dispatcher for free.
A refused request also emitted only a Content-length line, since the status
line for that branch was behind _DEBUG. Any client reads that as a protocol
error, which is how test 68 failed rather than reporting the refusal. Send a
403 instead.
Tests 68 and 77 posted without an id, which is what the engine used to accept,
so both now fetch the one the server renders into the form. Test 78 covers
accept, missing, empty and wrong, asserts the 403, and probes the key store
through ${projname} rather than the suppressed reply -- a reply-only assertion
passes even when the write goes through.
Also fix a leak that hung macOS CI: start() runs inside a command
substitution, so its $! never reached the parent and stop() guarded on an
empty variable, leaving one htsserver per call. Test 77 starts four, which is
exactly the four orphans the runner reported while sitting for half an hour
behind a green test log. Take the pid from the PID= line the server already
announces.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* webhttrack: bound the argv vector and escape quotes in the wizard command line
The wizard hands its httrack command line to the engine as one string, which
back_launch_cmd() split back into argv. Two things were wrong with that split.
It wrote into a fixed 1024-pointer vector with no bound, and every unquoted
space in the posted string yields an entry, so an ordinary mirror with a few
hundred URLs walked off the allocation. The split now lives in htscmdline.c as
hts_split_cmdline(), which sizes the vector from the separator count before
filling it, and the engine self-tests can reach it.
Quotes were also purely advisory: they toggled the "inside an argument" state
but nothing escaped them, so a double quote typed into a wizard field (user
agent, footer, path, project name) closed the argument early and the rest of
the value was parsed as fresh options -- among them -V, which reaches system().
Escaping has to happen where the argument boundary is known, so the template
gets its own ${arg:} filter for that context; ${html:} keeps its meaning for
the HTML attributes it is used in everywhere else, and HTML escaping would not
help anyway since the browser undoes it when it posts the command line back.
${arg:} backslash-escapes a quote and a backslash, and the splitter reads those
inside a quoted run, the same convention next_token() already implements for
doit.log. A value containing a quote now survives it intact instead of turning
into options.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* webhttrack: keep a stray quote in an unquoted field out of the split
The url and wildcard-filter fields go into the command line outside quotes,
where no backslash can escape anything: a single quote there flips the parity
of every quote after it, so a later escaped value ends up split as flags and
the escaping buys nothing. Emit %22 for those fields instead.
NULL-terminate the argv vector while here, matching the convention the tree
documents in htscharset.c, and fold the third copy of the entity table into
one helper.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: do not feed snprintf's return value back as its size argument
snprintf returns the length it wanted to write, so accumulating it blind
lets the next size argument wrap. The buffer is sized well past what the
loop needs, but the pattern is the one the project forbids.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Symbolize fatal-signal backtraces through addr2line
backtrace_symbols_fd() resolves names from .dynsym only, and
-fvisibility=hidden keeps every engine frame out of it, so a crash report
arrived as a column of bare module+offset. The handler now emits the raw trace
first and unconditionally, then groups the frames per module and runs addr2line
(or llvm-symbolizer) over the offsets, which reads DWARF and names the static
frames plus their inline chain.
-rdynamic is dropped: it only populated .dynsym and bought exactly one named
frame. -Wl,--build-id replaces it, so a trace from a stripped build can be
matched to its debug symbols.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Move the crash backtrace printer into src/htsbacktrace.c
httrack.c keeps only the two call sites. The symbolizer needs _GNU_SOURCE for
dladdr(), which is now confined to its own translation unit instead of being
forced on the whole CLI front-end.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Test: match glibc's backtrace format with or without the space
backtrace_symbols_fd() prints the trailing "[0xADDR]" with a leading space on
some glibc versions and without on others (Ubuntu 24.04), so the raw-frame
assertion failed everywhere but the dev box. Match only up to the offset.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Skip pseudo-modules with no file on disk
A frame in linux-vdso.so.1 made addr2line complain instead of resolving, so
the arm64 leg lost its symbolized output entirely. Renumber the test too:
77 landed on master with #700.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Stop discarding local symbols, which made the names wrong
--discard-all drops the local symbol of every static function, so addr2line
attributes the frame to the nearest surviving global: the abort frame read
dns_timeout_selftests instead of abortf_, with the file and line still right.
A wrong name is worse than none, and because the symbols go at link time no
-dbgsym package can recover them. Costs 21784 bytes on libhttrack.so.3, 0.6%.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Symbolize once when the handler itself faults
A fault inside the handler re-enters the printer, which interleaved a second
symbolized trace on the same fd and spent a second budget: 3.04s and two
overlapping reports, measured.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <xroche@gmail.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Drop NULL tests on inline array members
`lien_back::url_sav`, `htsblk::msg` and POSIX `dirent::d_name` are arrays,
so testing their address folds to a constant and gcc/clang report it
(-Waddress, -Wpointer-bool-conversion). Every site keeps whatever real
condition sat beside the dead one, so behavior is unchanged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* help_wizard: check the allocation, not the arrays it contains
The out-of-memory guard has been constant-false since d593418 folded the
nine separate wizard buffers into one struct: the names it tests are now
inline arrays, so `malloct()`'s result is never checked and an exhausted
heap gets a NULL-page write instead of the intended message. Also switch
the raw free() to freet() and release the struct on the two early returns
that leaked it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* proxytrack: test the WebDAV header fields, not their addresses
`PT_Element::lastmodified` and `::contenttype` are inline arrays, so both
`if`s were constant-true (-Waddress); use the `[0]` form the same file
already uses when it emits the GET headers. Neither is observable:
get_time_rfc822("") returns 0 and falls through to the index timestamp,
and proxytrack_add_DAV_Item already substitutes application/octet-stream
for an empty mime, which a PROPFIND probe against a cache entry carrying
no Content-Type confirms both before and after.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Fix the discarded const qualifiers rather than casting them away
`binput` and `cache_binput` only ever read through their source pointer,
so they take `const char *` now; that alone clears the cast in
htsrobots.c, and neither is exported nor declared in an installed header,
so no ABI question arises. `treathead` keeps `char *rcvd` because it does
NUL-cut the header in place, and the two selftest calls that fed it a
string literal get a mutable buffer instead, matching their three
siblings and removing a latent write to .rodata. The remaining two are
one-liners: zlib's `next_in` is already `const` under -DZLIB_CONST, and
htsback can call the non-const `jump_protocol` twin on its mutable
`url_adr`. libhttrack.vcxproj gains ZLIB_CONST so the MSVC build agrees
with autotools, as webhttrack and proxytrack already do.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: prove the WebDAV mime/timestamp fallback survives an empty field
proxytrack's DAV PROPFIND response computes a fallback content-type and
timestamp when a cache entry has no Content-Type/Last-Modified; that
fallback already existed before commit eae1dd0 changed the surrounding
always-true array-address checks, so this test guards the equivalence
rather than a bug. Verified it fails when the fallback default is
disabled, and passes unmodified against the pre-eae1dd0 code too.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: bound the PROPFIND request and cut the header comment
curl had no --max-time; an unbounded read wedges the runner instead of
failing it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: skip the WebDAV mime test on Windows
It is the first test to run proxytrack as a live listener, and MSYS cannot
reap a native one: the orphan wedged the whole Windows suite past its
45-minute budget, twice, destroying the log upload with it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* htsserver builds the redirect Location header in a 256-byte stack buffer
The POST redirect path checks strlen(file) but sprintf's newfile, which comes
straight from the client's "redirect" POST field with no length cap. A 300-byte
value overflows tmp[256]. The same value reached the Location header with no
CR/LF check, so it could split the response and inject headers.
Append into the dynamic String the other headers already use, and drop the
header entirely when the value carries a CR or LF.
The listen socket was SOCaddr_initany, so the server answered the LAN and not
just the local browser it exists to serve. Bind 127.0.0.1 by default, with
--bind <addr> to widen it again, resolved through the existing gethost() helper
the way proxytrack already does.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: satisfy shellcheck and shfmt in the new server test
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: drop the pre-fix narration from the oversized-value comment
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: assert the bound socket, not the announced URL
The listen-address assertions only compared the URL= banner, which is a
literal echo of argv: a build that announced 127.0.0.1 while binding the
wildcard passed. Probe 127.0.0.2 on the same port instead, which a wildcard
listener takes and a loopback-only one leaves free.
Also refuse an empty --bind, which fell through to every interface and
silently undid the new default.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: bound the response read and the empty --bind run
The recv() loop had no timeout and read until EOF; htsserver need not close
the connection after responding, which wedged the macOS runner for over an
hour. Stop at the end of the header block, which is all the test reads.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* htsserver: the session id must gate the request body, not the reply
Every field of a POST body is written straight into the one global key store
the templates and the command dispatcher both read, and "command" from there
reaches the engine. The gate ran after that write and compared "sid" against
"_sid" -- but "_sid" is copied into "sid" beforehand so the templates can
render it, so a request that simply omitted the field compared equal to
itself. Only a wrong id was refused; an absent one passed. Clearing the reply
afterwards does not help either, because the dispatcher sits outside the reply
guard.
Authenticate before parsing instead: scan the raw body for "sid", require at
least one occurrence and reject if any of them differs, and drop the body
untouched when it does not match. That leaves the shared template key alone,
and it closes the dispatcher for free.
A refused request also emitted only a Content-length line, since the status
line for that branch was behind _DEBUG. Any client reads that as a protocol
error, which is how test 68 failed rather than reporting the refusal. Send a
403 instead.
Tests 68 and 77 posted without an id, which is what the engine used to accept,
so both now fetch the one the server renders into the form. Test 78 covers
accept, missing, empty and wrong, asserts the 403, and probes the key store
through ${projname} rather than the suppressed reply -- a reply-only assertion
passes even when the write goes through.
Also fix a leak that hung macOS CI: start() runs inside a command
substitution, so its $! never reached the parent and stop() guarded on an
empty variable, leaving one htsserver per call. Test 77 starts four, which is
exactly the four orphans the runner reported while sitting for half an hour
behind a green test log. Take the pid from the PID= line the server already
announces.
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* -%S list file over 4GB overflows the heap
hts_main_internal() sized the -%S buffer as `cl + fz + 8192` and stored
the sum in an int url_sz, then fread() the untruncated 64-bit fz into it.
A 4GB+100KB rules file wraps the capacity to 110602 bytes on x64, the
realloct() succeeds, and the read walks off the heap. Intermediate sizes
land on a negative int and fail the allocation, which is luck, not design.
Route the file-size arithmetic through llint_grow_size_t(), a saturating
sibling of llint_to_size_t() that refuses a total it cannot represent, and
widen url_sz and the filelist offsets to size_t. The "config" sizing in the
same function and htscore.c's primary_len had the same shape: a file size
accumulated into an int before reaching an allocator. htscache.c's two
mirrored-file comparisons held a 64-bit fsize_utf8() in a size_t, which
truncates on Win32 and re-downloads a >4GB file already on disk.
Found by MSVC C4244 on x64; invisible to gcc/clang because int64_t to
size_t is width-preserving on LP64.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Print T_SOC with the format matching its width, not a bare %d
T_SOC is unsigned __int64 on Win64 (htsglobal.h): passing it to fprintf's
%d is undefined behavior, flagged by MSVC C4477. Add T_SOCP beside the
typedef, following the existing LLintP/INTsysP precedent, and use it at
both deletesoc() call sites (htslib.c:2601, :2607).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Fix MSVC C2099 in the new growsize self-test
A static const object used inside another object's static initializer is a
GNU/clang extension, not standard C: MSVC's /TC C mode rejects it ("initializer
is not a constant"). Replace the over32 local with a macro.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: cover the slack-only overrun and the largest capacity
A helper dropping the slack bound passed the table; -1 as extra also refused
either way, since llint_to_size_t() maps it to SIZE_MAX regardless.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* ARC cache replays a truncated HTTP reason phrase
proxytrack bounded the reason-phrase copy out of an ARC index by
sizeof(pos) - 1 where pos is a const char *, so a stored "404 Not Found"
replays as "404 Not Fou" (and "404 Not" on 32-bit). Use strncatbuff with
the destination's own size.
Fold the nine copies of the buff() family's source-capacity expression
into HTS_SIZEOF_SRC_, applying sizeof to the type so a decayed operand no
longer trips -Wsizeof-array-decay at five call sites; MSVC keeps the old
expression behind the guard HTS_IS_CHAR_BUFFER already uses. The two
other raw strncat calls become strncatbuff, htsbuff_catn stops handing
strnlen the (size_t)-1 sentinel, and htsweb.c no longer compares ep
against a NULL eps.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Size the new strsafe buffers away from sizeof(char*)
char[8] equals a pointer on LP64, so MSVC's array-vs-pointer heuristic
read the unterminated source as a pointer, skipped the bound and never
aborted; the x64 build failed while Win32 passed. Same trap the existing
comment in that function already warns about.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* configure: fix flag probes, move -rdynamic to the link line
AX_CHECK_COMPILE_FLAG only checks the exit status, and clang merely warns on
an unknown -W name, so -Wmissing-parameter-type reached every clang build and
warned on all 66 TUs. Probe with -Werror; -Wformat-nonliteral also needs
-Wformat there or gcc rejects it and the flag would be lost.
-rdynamic lived in DEFAULT_CFLAGS, i.e. AM_CPPFLAGS, so no link line ever saw
it; make it a link check. Drop -pie from CFLAGS_PIE (LDFLAGS_PIE has it), and
drop -Wdeclaration-after-statement, a C90 rule the gnu17 build does not follow
anywhere else.
Distinct build warnings: gcc 82 -> 54, clang 54 -> 41.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* configure: tighten the two new flag-block comments
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
WinHTTrack is an MBCS build, so it still lands on the ANSI-only Win32 entry points (TTN_NEEDTEXTA and the like) while the engine hands it UTF-8. The other direction is already exported as `hts_convertStringSystemToUTF8`, so this adds the mirror, `hts_convertStringUTF8ToSystem`, a one-liner over the existing `hts_convertStringCPFromUTF8`. The GUI can then delegate instead of keeping its own MultiByteToWideChar/WideCharToMultiByte copy (`CopyTextUTF8ToCP` in newlang.cpp, added while fixing #114). `hts_convertStringFromUTF8` would also have done the job, but it is declared plain `extern` and never reaches the DLL export table; left alone here. Windows-only addition, so no POSIX ABI change and no soname move.
The new `-#test=syscharset` self-test round-trips against the raw Win32 two-step it replaces, and skips off Windows.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* CLI: redraw the whole screen when the terminal is resized
The animated -%v display assumed a fixed 80x24 layout: it cleared the screen
once at start, and afterwards cleared only to end of line on the rows it
wrote. A resize left stale wrapped text around the frame, and on a terminal
shorter than the 20-row frame the stats block scrolled off the top at every
refresh.
Poll the terminal geometry at each refresh (TIOCGWINSZ, or the console screen
buffer info on Windows) and repaint in full when it moved. The in-progress
list is capped to the rows that fit, and the URL column follows the width; it
stays at the historical 40 characters on an 80-column terminal.
Closes#97
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* tests: assert the resize repaint, the row clamp and the URL width
The first pass only checked that a clear-screen followed one resize, which a
display clearing on every frame passes just as well, and neither the row clamp
nor the URL budget was exercised at all: the single 24x80 to 40x100 resize
leaves both at their 80x24 values.
The harness now waits for the display instead of sleeping, resizes width and
height separately, and asserts a repaint after each one, none in between, a
frame that fits a 10-row terminal, and a long URL that stops being truncated at
200 columns. The long path is a new /trickle/deep route, so the shared
/trickle/ index other tests assert on stays untouched. Mutating any of the four
behaviors away fails the test.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Dependabot only watched vcpkg, so the action pins drifted by hand:
windows-build.yml sat on actions/checkout@v4 while everything else moved to
v6. Add the github-actions ecosystem and bump that stray checkout to v6.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
vcpkg builds openssl/brotli/zlib/zstd from source on every Windows run.
Redirect its files binary cache into the workspace and persist it with
actions/cache, keyed on the manifest (which carries the builtin-baseline).
x-gha, which used to do this, was dropped from vcpkg-tool (#1662) after
GitHub changed the Actions cache API; actions/cache over the archive dir is
the endorsed replacement and stays within the GitHub-owned-actions policy.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Drop the removed Java .class parser from the credits
The Java binary .class parser was removed in #552, but its credit lingered:
"JavaParserClasses: Yann Philippot" in the WinHTTrack About box (lang.def plus
every lang/*.txt), and "for the java binary .class parser" in greetings.txt and
the contact page. Remove the About-box line and the now-defunct descriptor,
keeping Yann Philippot's name in the developed-by acknowledgments.
lang.def and lang/*.txt are read as strict line pairs, so the edit deletes the
substring inside the single credits line rather than removing any line, keeping
the pairing intact and each translated file's English msgid in lockstep with
lang.def. Verified by tests/62_lang-integrity.test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Also drop the localized java-parser credit from the translations
The About box shows the translation (msgstr), not the English msgid. The first
pass only removed the literal English "JavaParserClasses: Yann Philippot", so the
12 files whose translators localized the credit (Chinese, French, Russian,
Turkish, ...) still displayed it. Remove the localized credit line from each of
those msgstrs too; the pairing and every English msgid are untouched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Reword Yann Philippot's credit as a past contributor
Dropping the descriptor left a bare name in the developed-by list. Mark the
java binary .class parser as a past contribution instead, keeping the
acknowledgment meaningful now that the code is gone.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* contact.html: use a literal dot in .class
The " dot " spelling is the page's email anti-harvest obfuscation; applied to a
filename extension it just left a stray double space. The page already writes
literal dots elsewhere (v2.0, v3.0).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Bump the version in configure.ac, htsglobal.h and version.rc, move VERSION_INFO to 3:6:0, and add the history.txt and debian/changelog entries.
The WARC fields landed at the tail of httrackp and htsblk, but the embedded htsoptstate and htsblk members sit mid-struct, so the fields after them shifted. Installed layouts moved while the soname stays .so.3; the VERSION_INFO comment records that rather than claiming the layouts are unchanged.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Windows: import IE cookies via FindFirstFileW so long/non-ASCII jars load
cookie_load's IE-cookie scan and its cookies.txt read used the narrow
ANSI file API (FindFirstFileA, fopen, remove), so a cookie folder with a
non-ASCII or >MAX_PATH path was silently skipped, and any matched IE
cookie name was fed back through the mirror path as CP_ACP bytes. Route
the glob through hts_pathToUCS2 + FindFirstFileW, convert cFileName to
UTF-8 before rebuilding the path, and use the FOPEN/UNLINK wrappers.
The A->W find and the sink swaps land together on purpose: swapping only
the sink would feed a CP_ACP name into hts_fopen_utf8's UTF-8 decode and
mojibake the accidentally-consistent ANSI path. hts_pathToUCS2 is
un-static'd (internal, still -fvisibility=hidden / not HTSEXT_API) so the
glob gets the same \\?\ long-path treatment as the other wrappers.
Last piece of the #133 Windows long-path series. Self-test
-#test=cookieimport drives cookie_load against a long, non-ASCII folder;
POSIX is a positive control (IE block compiled out), the wide glob and
IE import run on the Windows legs.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Trim review-flagged comments to one line
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Windows: enumerate long or non-ASCII directories via FindFirstFileW
The opendir/readdir emulation was fully ANSI (FindFirstFileA, CP_ACP
cFileName), so it capped listing at MAX_PATH and mis-decoded a non-ASCII
path — the odd wrapper out in an engine that feeds UTF-8 paths and reads
d_name as UTF-8. The one live Windows caller, the end-of-crawl
hts-cache/ref cleanup, silently no-op'd at a long or non-ASCII project
root, leaking the temp ref/ directory into the finished mirror.
Route both through hts_pathToUCS2 (\\?\ prefixing, #684) and the wide
APIs, converting cFileName back to UTF-8. d_name now holds UTF-8, so
HTS_DIRENT_SIZE grows to 1024 to fit MAX_PATH's worst-case expansion; the
struct is _WIN32-only, no POSIX ABI impact. Self-test direnum drives it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Trim comments to one line each (review follow-up)
Condense the opendir/readdir and direnum-selftest comments per the
review's comment-conciseness pass; no behavior change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* topindex: convert the winprofile.ini category to UTF-8 on Windows (#216)
The project category read from hts-cache/winprofile.ini comes back in the
Windows ANSI codepage and is written into the charset=utf-8 topindex
template, so a non-ASCII category renders as mojibake. Convert it the same
way #681 did the project name. The topindex self-test now writes a
winprofile.ini with a non-ASCII category and asserts the generated index
carries the UTF-8 form.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* topindex: drop the redundant category-assert comment
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The 01_engine-fsize self-test builds a 5GB sparse file to exercise the 32-bit size wrap. On GNU/Hurd i386 the ext2fs translator rejects the extend with EFBIG, failing the Debian build on that port. Return 77 (automake skip) instead of 1 when the extend fails with EFBIG, and map that to a skip in the .test wrapper. Any other errno, or a wrong size report, still fails, so real regressions are not masked.
* Windows: route mirror-tree file ops through the UTF-8/long-path wrappers
The engine's raw fopen/remove/rename/rmdir/mkdir/fexist/fsize calls on
mirror-tree paths bypassed the UTF-8 (_w*) wrappers, so on Windows a mirror
under a non-ASCII or >MAX_PATH directory silently broke: a long/non-ASCII
file read as absent (its unlink/reget skipped), or its bytes went to a
mojibaked, truncated path.
Swap those call sites in htscoremain/htscore/htsparse/htsindex/htshelp to
the FOPEN/UNLINK/RENAME/MKDIR/fexist_utf8/fsize_utf8 twins, and add
hts_rmdir_utf8 (RMDIR) for the three rmdir sites that lacked a wrapper. The
arguments are already UTF-8 (Windows argv is transcoded at startup, fconcat
is byte-transparent), so there is no double-conversion. Raw calls on
genuinely system-charset or ASCII paths (getenv/$HOME, structcheck's ASCII
twin, "config") are left as-is.
-#test=mirrorio drives a long AND non-ASCII path through the wrapped guards
and the new rmdir wrapper; on POSIX it stands as a byte-transparent control.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* test: cover RENAME on the long+non-ASCII mirror path
The mirrorio self-test drove FOPEN/UNLINK/RMDIR but not RENAME, which the
sweep newly routes to. Rename the leaf to a non-ASCII sibling and verify
the move before teardown.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Windows: reach past MAX_PATH via \\?\ at the file-op choke-point (#133)
The _w* file wrappers on Windows all funnel their path through one
converter. Route it through a new hts_pathToUCS2 that, for a path near
MAX_PATH, absolutizes and normalizes it with GetFullPathNameW and
prepends the "\\?\" verbatim prefix ("\\?\UNC\" for UNC shares) so the
wide file APIs accept paths past 260 chars. Shorter paths keep today's
exact behavior, so nothing changes until the naming ceiling in htsname.c
is raised (a follow-up); a -#test=longpath self-test drives a >260-char
path through the wrappers to exercise the new branch on Windows CI now.
Partially addresses #133.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Run the long-path self-test on the Windows CI leg; tighten comments
The Windows job iterates an explicit test glob (01_engine-*, 01_zlib-*,
*_local-*, ...), so 76_engine-longpath-io.test matched nothing and the
\\?\ path never ran on the one platform that needs it. Rename it to
01_engine-longpath-io.test so the glob picks it up and the prefixing is
actually exercised on Windows. Also trim the review comments to the why.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Use the platform path limit off Windows, not the Windows MAX_PATH (#133)
url_savename capped every saved path at 236 chars (MAX_PATH minus 8.3
headroom minus the ".delayed" marker) on all platforms, so Linux, macOS
and Android hashed long names to fit a limit only Windows has. Off
Windows, derive the ceiling from the platform's own PATH_MAX/NAME_MAX,
clamped to the fixed save buffer so an oversized output dir can never
push the final path past it and abort. Windows keeps the MAX_PATH
ceiling until the engine can prefix its paths with \\?\.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Cut the save name in bytes, not codepoints, so multibyte paths can't overflow
The path-length guard measured length in UTF-8 codepoints (hts_stringLengthUTF8)
while the buffer that receives parent+name is a fixed byte array whose overflow
aborts() rather than truncates. With the ceiling raised to ~1984 on POSIX, a
multibyte name (e.g. a long CJK path) can stay under the codepoint cap yet run
its byte length past the 2048-byte buffer, crashing the crawl on the final
prepend. Add an unconditional byte-boundary cut before the parent is prepended,
and drop the earlier codepoint-unit clamp it supersedes. Covers the oversized
-O parent case in the same measure. Regression test uses a 339-char CJK name
under a ~1020-byte output dir, which aborted before the cut.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Leave an oversized parent to the existing abort, not an empty colliding name
When the output dir alone fills the save buffer, forcing the name to empty made
every URL under it collapse to the same path, feeding the unbounded collision
sprintf and turning a clean abort into a 1-2 byte overflow. Only shrink the
name; a buffer-filling parent aborts on the prepend exactly as before (that
oversized-parent path is pre-existing and unrelated to the #133 ceiling).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* topindex: non-ASCII project names render as mojibake on Windows (#216)
hts_buildtopindex() lists each sub-project by the name FindFirstFileA
returns in the ANSI codepage, then writes it into a document declaring
charset=utf-8, so a non-ASCII name shows up as mojibake. Convert the name
to UTF-8 on Windows, matching the gif-path fix already in this function
for #217. The topindex self-test now builds a non-ASCII sub-project and
asserts the generated index carries the UTF-8 name.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* topindex: drop const so freet() can null the converted name (MSVC)
The #ifdef _WIN32 conversion block only compiles on Windows, where MSVC
rejected freet() nulling a char *const (C2166). Make the pointer mutable.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* topindex: tighten the added comments
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The command-line guide's recipe section had no entry for WARC output, so
the five --warc* flags were discoverable only through --help or the man
page. Add a §11 recipe leading with the gotcha that trips people: the
archive is written alongside the browsable mirror, not instead of it.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
WARC is a transaction-level archive of what was fetched (a sibling of the
hts-cache), not a browsable-mirror build format like MHTML, so it belongs in
the Log/index/cache group next to the cache options. Help text only; the
webhttrack GUI already places it on the Log/Index/Cache tab (option9).
Signed-off-by: Xavier Roche <xroche@gmail.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Verbatim compressed bodies (Content-Encoding kept, only hop-by-hop
Transfer-Encoding dropped, Content-Length rewritten to the stored length)
are now the sole behavior, matching wget --warc and Heritrix. The former
default that decoded the body and stripped Content-Encoding is gone, along
with the --warc-verbatim switch (added the same day, never released), the
warc_verbatim option field, and the dual-mode plumbing in
normalize_http_headers.
This also fixes a cap-truncated compressed response: it previously stored
the raw compressed partial while stripping Content-Encoding, mislabeling
gzip bytes as identity. Keeping Content-Encoding makes the record's label
match its body; the truncation still carries WARC-Truncated: length/time.
header_is now tolerates whitespace before the ':' so a non-compliant
"Content-Encoding : gzip" is still recognized.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Add --warc-cdx: sorted CDXJ index alongside the WARC archive
WARC v2 PR B. --warc-cdx (alias --warc-cdxj, sub-flag -%rc) writes a sorted
CDXJ index next to the .warc.gz, one line per response/revisit/resource record
(warcinfo/request are not indexed): the SURT sort key, a 14-digit timestamp,
and a JSON object with url, mime, status, payload digest, and the record's byte
offset and length in the gzip stream so a replay tool can seek and inflate a
single member. Rotation-aware (the filename field tracks the current segment).
SURT canonicalization is new (surt_canon in htswarc.c): scheme/userinfo
dropped, host lowercased with a leading www[digits] label and the scheme
default port stripped, labels reversed and comma-joined, a non-default port
kept, IPv4 and [IPv6] literals left verbatim. Lines are accumulated in the
writer and qsort'd in LC_ALL=C byte order at close.
Tested by -#test=warc-surt (SURT vectors, 01_engine, MSan-instrumented) and
-#test=warc-cdx (end-to-end: sorted, one line per record, each offset/length
independently inflates to the matching WARC-Target-URI; 01_zlib).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Add --wacz: package the WARC archive, CDXJ index and pages as a WACZ file
--wacz bundles the crawl's WARC segment(s), the sorted CDXJ index and a
generated pages.jsonl into a single WACZ 1.1.1 package at crawl end, using
the in-tree minizip. It implies --warc and --warc-cdx. Every ZIP entry is
stored (no re-deflate of the already-gzipped WARC), as the WACZ spec
requires. datapackage.json lists each file with its SHA-256 and size, and
datapackage-digest.json chains the digest of datapackage.json.
SHA-256 comes from OpenSSL; a build without OpenSSL cannot emit conformant
digests, so --wacz is refused there with a clear log line while the WARC and
CDXJ are still written.
pages.jsonl captures the top-level 200 text/html responses (URL + WARC-Date),
bounded like the CDX accumulator; the first seeds datapackage mainPageUrl.
Tests: an in-process self-test (-#test=warc-wacz) unzips the package and
asserts the layout, STORE mode on every entry, each recomputed SHA-256, the
digest chain, and the pages header; a crawl test validates a real .wacz with
a stdlib validator (py-wacz/pywb when importable). Both skip cleanly without
OpenSSL.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Package WACZ atomically so a failed re-run can't destroy a good archive
warc_wacz_package opened <base>.wacz directly with a create/truncate zip
handle, wiping the previous archive before it checked its inputs existed. A
re-run that produced zero indexable records (empty crawl, or every fetch
erroring so warc_cdx_flush writes no .cdx) truncated the good .wacz, failed to
find indexes/index.cdx, then unlinked the now-empty file with no log. Same
data-loss class as the hard-abort file truncation (#522).
Build the package into <base>.wacz.tmp and only rename it over <base>.wacz on
full success (RENAME, with an unlink+rename fallback so Windows can clobber).
On any error unlink the temp, leave the existing archive untouched, and warn
(the old error path was silent); the writer keeps opt for close-time logging.
The warc-wacz self-test now proves it: after a good package it drops the .cdx,
re-runs empty, and asserts the .wacz is byte-unchanged. Also tightens both the
C self-test and the Python validator against the WACZ spec (profile,
wacz_version, digest path, and >= 1 pages.jsonl body row with url + ts) so a
writer emitting an empty package can't pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* WARC: add --warc-verbatim to store compressed bodies as received
--warc-verbatim (internal -%rv, implies --warc) captures a compressed
response in its as-received content-coded form instead of the decoded
strategy-B body. back_finalize already materializes the whole de-chunked
compressed body as a temp file just before decoding it, so the tee is a
save-before-unlink: adopt that spool onto the new htsblk.warc_rawpath /
warc_rawsize instead of unlinking it, and emit the response record with
Content-Encoding kept and Content-Length set to the compressed length
(normalize_http_headers gains a keep_ce mode). WARC-Payload-Digest is
then over the coded payload, which is what the record carries.
Non-compressed responses, the never-spooled is_write archive-ext case,
and the cap-truncated second emit site all leave warc_rawpath NULL and
fall back to strategy B unchanged. The spool is owned by the entry:
freed and unlinked in warc_free_request, NULLed in back_copy_static and
back_unserialize so a shallow copy never double-unlinks.
Output is now byte-replay-identical to what the origin sent, matching
wget --warc and Heritrix.
Tests: a warc-verbatim engine self-test asserts the response keeps
Content-Encoding: gzip, Content-Length == the compressed length, the
stored bytes equal the gzip input and inflate to the known plaintext
(bite-checked); tests/74_local-warc-verbatim crawls a gzip-served
fixture and runs the strategy-A/B differential through
warc-validate.py --verbatim (inflate(stored) == the decoded body).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* WARC verbatim: cover the direct-to-disk adoption path, harden the validator
Review follow-up to --warc-verbatim (strategy A).
The spool adoption in back_finalize runs on both the in-memory and the
direct-to-disk (is_write) branch, but only the in-memory path was crawl-tested.
Add a non-hypertext gzip fixture (application/octet-stream) so the crawler
streams it to disk and the is_write adoption is exercised; the warc-validate
--verbatim gate now asserts Content-Encoding kept, Content-Length == compressed
length, and inflate(stored) == the served plaintext for both assets. Confirmed
at runtime that page.html takes is_write=0 and data.bin is_write=1, and
bite-checked that dropping the is_write adoption fails the new assertion.
warc-validate --verbatim now requires WARC-Payload-Digest on a verbatim record
when the file emits digests at all (an OpenSSL build), so a regression that
drops the digest is caught without breaking the no-OpenSSL leg.
Extract warc_adopt_rawspool() so the WARC spool bookkeeping lives behind the
htswarc boundary, mirroring warc_stash_response. Comment trims throughout.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
WARC v2 PR B. --warc-cdx (alias --warc-cdxj, sub-flag -%rc) writes a sorted
CDXJ index next to the .warc.gz, one line per response/revisit/resource record
(warcinfo/request are not indexed): the SURT sort key, a 14-digit timestamp,
and a JSON object with url, mime, status, payload digest, and the record's byte
offset and length in the gzip stream so a replay tool can seek and inflate a
single member. Rotation-aware (the filename field tracks the current segment).
SURT canonicalization is new (surt_canon in htswarc.c): scheme/userinfo
dropped, host lowercased with a leading www[digits] label and the scheme
default port stripped, labels reversed and comma-joined, a non-default port
kept, IPv4 and [IPv6] literals left verbatim. Lines are accumulated in the
writer and qsort'd in LC_ALL=C byte order at close.
Tested by -#test=warc-surt (SURT vectors, 01_engine, MSan-instrumented) and
-#test=warc-cdx (end-to-end: sorted, one line per record, each offset/length
independently inflates to the matching WARC-Target-URI; 01_zlib).
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Expose --warc / --warc-file in the webhttrack GUI
Add a WARC output checkbox and an optional archive-name field to the
"Log, Index, Cache" option tab, beside store-all-in-cache. The checkbox
emits --warc (auto-named archive); the text field emits --warc-file NAME.
Wiring mirrors how #589 added cookies-file and strip-query: the option9.html
form fields, the generated httrack command and winprofile.ini in step4.html,
and the reload remap in step2.html.
New LANG_WARC / LANG_WARCFILE strings and tooltips land in lang.def,
English.txt and Francais.txt; the remaining 28 language files fall back to
French for now and are a follow-up.
The webhttrack smoke test now also fetches option9.html and requires the
WARC control to render.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Translate the webhttrack WARC strings into the remaining languages
PR #672 added a WARC toggle to the webhttrack GUI with 4 new LANG keys,
translated only in English and French. Append the same 4 msgid/translation
pairs to the other 28 lang/*.txt so those locales stop falling back to
English. Each translation is encoded in the file's declared LANGUAGE_CHARSET
(the charset the server serves the page and parses the form as) with a strict
lossless round-trip, matching each file's existing CRLF/LF line ending.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Add WARC/1.1 v1.1: --warc-max-size rotation, WARC-Truncated, FTP resource records
Extends the merged WARC writer with the three v1.1 items. All logic stays in
htswarc.c; the engine hooks stay thin.
--warc-max-size N (-%rs) rotates the archive into NAME-00000.warc.gz, -00001,
... once a segment passes N bytes (wget naming), each segment led by its own
warcinfo and never splitting a record. N<=0 keeps the single-file behavior.
WARC-Truncated tags a body cut short by a cap. The mirror size (-M) and time
(-E) caps abort in-flight transfers and overwrite the slot's status to a
negative TIMEOUT, after which back_finalize and the WARC hook both bail; so the
partial is archived at the abort site, before the clobber, with WARC-Truncated:
length/time. HTTrack's own incomplete/retry bookkeeping is untouched, and a
genuine broken transfer (not a cap) is still not archived. The per-file cap (-m)
rejects on Content-Length rather than truncating, so it has no truncated body.
Standard ISO 28500 tokens only (length/time/disconnect); the task's non-standard
"disk" is left out.
FTP transfers have no HTTP envelope, so an ftp:// capture becomes one resource
record: WARC-Type: resource, the payload's own Content-Type, block = payload,
no request/response. FTP back_finalize runs on the main loop (the worker thread
only flips status), so the single-writer no-lock design holds.
Self-tests warc-trunc / warc-ftp / warc-rotate added to the -#test=warc family
and 01_zlib-warc.test; each was confirmed to fail without the corresponding
code. man/httrack.1 + html regenerated for the %r help line.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Range-check --warc-max-size and regenerate the man page
Replace the unchecked sscanf on the --warc-max-size argument with a
strtoll parse that rejects non-numeric, negative, and overflowing values,
leaving the default 0 (single archive) intact. The man page and its HTML
render were already regenerated for the option's help line in the feature
commit, so this only tightens the parse.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds `--warc` / `--warc-file NAME`, writing an ISO-28500 WARC/1.1 archive of a crawl (warcinfo + request + response records, gzip per record, SHA-1 digests under OpenSSL) so HTTrack output replays in Wayback-style tools (ReplayWeb.page, pywb). Under `--update`, unchanged resources become `revisit` records instead of duplicate copies. No new dependency.
Fidelity note: HTTrack decodes gzip/br/zstd inline, so v1 stores the decoded body and normalizes the encoding headers (drops `Content-Encoding`/`Transfer-Encoding`, recomputes `Content-Length`). Valid and replayable; verbatim-compressed capture and WACZ are v2. Logic is isolated in a new `src/warc.c`.
First phase (v1) of #668.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Skip an oversized -%F footer instead of aborting the crawl
The per-page footer is expanded into a fixed ~3KB stack buffer. On overflow
hts_footer_format returns <0 and leaves the buffer unterminated, but the call
site ignored the return and ran strcatbuff on it regardless; strcatbuff's
bounded strlen finds no terminator within capacity and abort()s, killing the
crawl with SIGABRT. Reachable whenever the footer expansion exceeds the buffer
(a field referenced several times with a long URL/path, or a URL that triples
under the HTML-comment escaping); the pre-existing default footer has the same
exposure.
Guard the emit on the formatter's return: on overflow, drop the footer rather
than crash. Regression test drives a file:// crawl of a deep path with a footer
repeating {path} and asserts the crawl completes; it aborts on the pre-fix
binary.
Closes#669
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Skip the footer-overflow test on Windows (MAX_PATH)
The test triggers the overflow with a path longer than Windows MAX_PATH (260),
which the source tree and mirror output both exceed there, so httrack can't
create it and the crawl fails for an unrelated reason. Restrict it to POSIX; the
fix is platform-independent and the formatter's overflow-return contract is
still covered cross-platform by the footerfmt self-test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Allowlist the Windows skip of the footer-overflow test
The Windows CI pins expected skips so an all-skipped suite can't pass green;
01_engine-footer-overflow.test now skips there (MAX_PATH), so add it to the list.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Add named {addr}{path}{date}{version} footer fields alongside the legacy %s form
The -%F footer was a bare positional printf: each %s consumed the next of a
fixed addr/path/date/version arg list. It is order-coupled and inexpressive
(you cannot reach {date} without emitting addr and path first, cannot reorder,
and a wrong %s count silently yields "???" or shifted values), and the help
text advertised a cryptic bracket syntax.
hts_footer_format now dispatches on content: a footer containing %s keeps the
legacy positional model byte-for-byte (the default footer and every existing
-%F string are unaffected), while a footer without %s uses named fields
{addr} {path} {date} {version}, with "{{"/"}}" for a literal brace and any
unrecognized {...} left verbatim so typos stay visible. Sanitization is
unchanged: addr/path still pass through html_inline_safe() at the call site,
so the comment-injection guard (#165) still holds.
Driven by a new -#test=footerfmt self-test (tests/01_engine-footerfmt.test)
covering both models, brace escaping, the mixed-mode dispatch boundary, and
the overflow/zero-size return paths.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Regenerate html/httrack.man.html for the -%F help change
The man/html sync CI guard requires html/httrack.man.html to change whenever
man/httrack.1 does; the -%F help-text update regenerated the man page but not
its html rendering.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Add {url}{lastmodified}{mime}{charset}{status}{size} footer fields
Extends the named footer set beyond addr/path/date/version. hts_footer_format
now takes a {name,value} table instead of four fixed parameters, so the field
set is data rather than signature and the legacy %s path looks addr/path/date/
version up by name (order-independent, still byte-for-byte). The call site
supplies the new fields from the current response: {url} (scheme + host + path,
credentials stripped like {addr}), {lastmodified}, {mime}, {charset}, {status}
and {size}.
Every network-derived string is html_inline_safe()'d, since the footer sits
inside an HTML comment and a value holding "-->" would otherwise close it and
inject markup (#165); {status}/{size} are formatted integers and need none.
Documented in --help, the man pages and the Command-Line Guide, and covered by
the -#test=footerfmt self-test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The offline crawl harness reads the server's ephemeral port from the "PORT <n>"
line it prints once bound, but only waited 5s (50 x 0.1s) and matched head -n1.
Under `make check -jN` up to 16 Python servers cold-start at once; on a loaded
Windows runner a cold MSYS Python start can lag past 5s, so discovery timed out
with an empty log ("could not discover server port:") and failed 13_local-cookies
spuriously. Give it a 30s deadline, and match the PORT line anywhere so a stray
startup warning merged via 2>&1 can't wedge discovery until timeout.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
#R was listed twice: "cache repair" and "old FTP routines". The FTP handler
is commented out in the parser, so the only live meaning is cache repair;
drop the bogus line. #X is a no-op (the parser prints "option has no effect"),
so remove its help line and misleading default marker.
Add help lines for three real but undocumented options: -%z
(--disable-compression), -%t (keep the original file extension), and -y
(--background-on-suspend).
Regenerate man/httrack.1 and html/httrack.man.html from the updated --help.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The offline suite waited on each httrack crawl with a bare wait bounded only by the engine's --max-time; a fetch that wedges past that (a Windows socket stall the engine misses) blocked wait forever and ran the test step to its 45-minute cap. wait_bounded attaches the #595 kill_tree reaper to the crawl pid so an overrun is reaped in seconds, stop_server now reaps the server's native tree, and 72_watchdog-crawl proves the fail-fast against an always-stall endpoint.
The "Select URLs" and "Start" section headings on the webhttrack server
wizard were literal English, bypassing the ${LANG_...} substitution, so
they stayed English in all 30 locales. Reuse existing, already-translated
keys: step3 -> ${LANG_G44} ("Web Addresses: (URL)"), step4 -> ${LANG_J9}
("Start"). No lang/*.txt changes needed.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Correct spelling errors (beginning, dishonest, redistributing, forbid,
personal, chosen, address, occurrences) and refresh stale content: collapse
the obsolete per-Windows-version questions into a single current statement,
drop dead OS keywords, describe the cookies.txt / --cookies-file workflow
instead of the Netscape/IE folder steps, and reword the rtsp note to say
streaming protocols are out of scope rather than "not supported yet".
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Correct several spelling typos in html/abuse.html and html/filters.html,
and point filters.html readers at the --why (-%Y) option, which reports
which filter rule accepted or blocked a given URL.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The 2007 cmddoc.html beginner walkthrough is fully subsumed by cmdguide.html.
Repoint index.html's "Command-line version" bullet to httrack.man.html (the
generated reference, previously unlinked) and mark the now-orphaned legacy
options.html as possibly-stale, pointing at the man page and the guide. Also
fixes two long-standing index.html typos (relese, Developper).
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Fix the download-PDFs example in the command-line guide
The old example ran httrack with no filters, so it mirrored the whole
site rather than downloading its PDFs. The corrected command restricts
to the site, admits the HTML pages as scaffolding for link discovery
(including directory-index pages via *[file]/ so sub-directory PDFs are
found), keeps the PDFs, and drops everything else.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Use *[path]/ so the PDF recipe follows nested directory indexes
*[file] cannot cross a slash, so *[file]/ only admits a single directory level and
silently drops PDFs linked under nested paths (dated archives, doc trees). *[path]/
spans slashes and follows directory-index pages at any depth, with no over-admission
(only trailing-slash URLs match, never files).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
html/httrack.man.html is groff-rendered from man/httrack.1 and committed, but
had drifted far behind the roff page (missing the option table-of-contents and
many options such as --pause and --delayed-type-check). Regenerate it, and make
regen-man-html strip groff's version-stamp and creation-date comments so the
committed file no longer churns across groff versions or rebuilds.
Add a lightweight CI guard: a PR that changes man/httrack.1 must also change
html/httrack.man.html, catching the "regenerated the roff page, forgot the HTML"
case that caused this drift. Rendering needs the full groff html device, so CI
verifies the two move together rather than regenerating in-CI.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add a "Filter wildcards" subsection listing the bracket wildcard forms
the matcher (src/htsfilters.c) implements: *, *[file]/*[name], *[path],
*[param], character sets and ranges, literal escapes, size rules, and
the *[] end anchor. It links to the full filters page for the rest.
Added as an h4 under the filters section so section numbering and the
page ToC are untouched.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
src/minizip/ vendors patched minizip as X.orig + X.diff pairs; X.diff is
replayed onto fresh upstream on the next re-sync, so a live edit that skips
X.diff silently reverts then. PR #640 fixed a signed-shift UB in the live
mztools.c (READ_32 now casts each recovered 16-bit half to uLong before the
shift) but never regenerated mztools.c.diff, so a re-sync would drop it.
Regenerate mztools.c.diff from the live file; it reproduces mztools.c
byte-for-byte again. Header date labels keep the existing convention so the
change is content-only. The other four snapshots already match and are untouched.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Non-ASCII single -O drops the hts-cache into a mangled twin directory on Windows
PR #636 fixed the log half of #630 and left the cache for a follow-up. With a
single -O café, path_log holds UTF-8 bytes, but cache_init created hts-cache and
opened new.zip through ANSI calls, so on Windows the cache landed in a second
mangled café-twin directory beside the mirror, and a later --update read the
cache from the wrong place.
This routes every path_log filesystem op in cache_init and the reconcile helpers
through the UTF-8 wrappers (structcheck_utf8, MKDIR, FOPEN, UNLINK, RENAME,
fexist_utf8, fsize_utf8), and opens the cache ZIP through minizip's filefunc
entry points (zipOpen2/unzOpen2) backed by hts_fopen_utf8. On POSIX the wrappers
resolve to the same libc calls, so only Windows changes. The conversion is
all-or-nothing: a directory created UTF-8 but a ZIP opened ANSI would split the
cache, so the whole cache-open path flips together.
The corrupt-cache repair path stays ANSI. mztools' unzRepair takes no filefunc,
so a corrupt cache under a non-ASCII path fails to repair (the site is
re-crawled) rather than forking a twin.
Test 69 now also asserts the cache lands under the café directory through a new
--cache-under-logroot audit. Like the log assertion it only bites on the Windows
CI leg; on POSIX it exercises the code path without reproducing the mojibake.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Cache UTF-8 filefunc must keep 64-bit offsets and the cache dir 0700
The #630 cache-UTF-8 hook built a 32-bit zlib_filefunc_def via
fill_fopen_filefunc and called unzOpen2/zipOpen2. Minizip runs a 32-bit
filefunc through fill_zlib_filefunc64_32_def_from_filefunc32, which NULLs the
64-bit seek so every seek truncates the offset to uLong: on Windows LLP64 a
cache >=4GB hard-fails and >=2GB corrupts. The ANSI code this replaced went
through fill_fopen64_filefunc and never truncated. Keep 64-bit: fill the
zlib_filefunc64_def and override only zopen64_file with the UTF-8 opener, then
call unzOpen2_64/zipOpen2_64.
Also restore the POSIX cache-dir mode: the MKDIR macro creates with
HTS_ACCESS_FOLDER (0755), loosening the 0700 the explicit mkdir used. There is
no UTF-8 mkdir-with-mode, so keep the platform split - UTF-8 MKDIR on Windows
(which ignores the mode), mkdir(HTS_PROTECT_FOLDER) on POSIX.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Tighten the #630 comments
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Strip only the scheme's own default port, not :80 on every scheme
hts_strip_default_port treated 80 as the default for every scheme, so an
explicit :80 on a non-http URL was dropped and the fetch silently moved to that
scheme's real default: https://h:80/x became https://h/x (fetched on 443), ftp
likewise. The reverse gap: a scheme's own default (443 https, 21 ftp) was never
stripped, so https://h:443/x kept a redundant port and would not dedup against
the portless form.
Derive the default from lien's scheme (80 http, 443 https, 21 ftp; 80 when
absent or unknown) and strip only when the port equals that. Same family as
#627/#614.
Closes#638
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Exclude the auth-bypass CodeQL query and pin case-insensitive scheme matching
cpp/user-controlled-bypass models auth-bypass-by-spoofing, but httrack has no
auth or access-control surface; its nearest downstream decision, the +/- crawl
filter, is a mirror boundary, not a security one. Exclude it like the
world-writable query.
The stripport self-test's scheme detection is case-insensitive (strfield/streql)
but nothing pinned it: a case-sensitive rewrite would strip :80 from HTTPS://
unnoticed. Replace the non-discriminating :8443 case (no scheme defaults to 8443,
so it passes for any impl) with HTTPS://h:80, which must keep its port. Note at
scheme_default_port that schemeless/protocol-relative links default to 80.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Tighten the #638 comments
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Force a whole-file refetch when a rejected 206 resume can loop (#581)
Resuming an interrupted download, the engine sends a Range and, if the
server answers an unusable 206 (Content-Range not matching the partial on
disk), drops the partial and retries the whole file. On Windows a hard-killed
first pass leaves a truncated cache that the next pass must repair; the retry
can then re-derive a Range from a partial or temp-ref that outlived the
restart, meet the same unusable 206, and loop until retries run out. The
partial is removed but never refetched, so the file is lost from the mirror.
Carry a refetch-whole signal from the restart-whole decision onto the
requeued link, so the retry's back_add drops any stale temp-ref and skips the
partial/temp-ref resume branches: it sends no Range and GETs the whole file,
regardless of whether a partial survived the restart. On POSIX the restart
already removes both sources, so the retry was already Range-less and behavior
is unchanged; test 71 still recovers the file whole.
back_add gains an internal (hidden, non-exported) parameter; htsblk and
lien_url gain one trailing field each. No exported symbol changes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Use hts_boolean and HTS_TRUE/HTS_FALSE for the refetch-whole flag
The #581 fix carried the whole-file refetch flag under three inconsistent
types (short int, char, int) set with bare 1/0. Normalize all three to the
house hts_boolean type and use the HTS_TRUE/HTS_FALSE macros for the literal
sets. No behavior change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Advertise -%N's long option and document the real long forms
The help display appends each option's long alias by looking it up in the
alias table, but it first stripped a trailing N to turn placeholders like
cN into c. That also turned -%N into -%, so --delayed-type-check was never
advertised in --help (nor in the generated man page). Try the flag as-is
first and only strip a trailing N on a miss; -%N now shows its long form,
and cN still resolves to --sockets.
The command-line guide had several rows marked short-only that in fact
have long aliases (I had read them off a stale installed binary): fill in
--pause, --strip-query, --disable-compression, the three --keep-* dedup
opts, and --delayed-type-check. Regenerate man/httrack.1 to match, and
drop the long-dead commented-out -%O chroot help line.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Bound the option-token sscanf to the buffer size
CodeQL flagged the %s read into cmd[32] as an unbounded copy on the line
the previous commit touched. The tokens come from HTTrack's own help
strings, not hostile input, so it was not reachable in practice, but the
copy should be bounded regardless. Limit it to %30s (the buffer holds the
leading '-' plus 30 chars and a NUL).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Add a task-oriented command-line guide to the offline docs
The command-line docs so far are the option list and the generated manual
page: exhaustive, but organized by flag, not by task. Newcomers arrive
expecting wget/curl syntax and hit the same walls (only the index came
down, filters that do the opposite of what they read, an update that
deletes files), because the reference answers "what does -X do", not "how
do I do the thing I want".
cmdguide.html is a task-oriented layer on top of the manual page: quick
start, scope, filters, limits, naming, links, identity/login, proxy,
update/cache, scripting, then eleven copy-ready recipes with the one
gotcha each. It foregrounds the defaults that actually surprise people
(the ~100 KB/s rate cap, the -c/-A/-%c security clamps, --update purge
semantics, the depth off-by-one, mime filters running after headers) and
links the manual page for per-option detail. Linked from cmddoc.html and
the index. Every documented default and behavior was checked against the
engine source.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Correct four scope/option descriptions in the guide
Adversarial review against the engine source caught four mislabels:
-d is "same principal domain" and -l is "same TLD" (not "same directory"
and "same domain"); -U goes up only and -B goes both ways (the guide had
-U as up-and-down and -B as "anywhere"); -%M archives the whole mirror
into one index.mht, not each page; and -t HEAD-tests links outside scope
rather than reporting "what a scope would reach". The rest of the guide's
documented behavior verified against code.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Fix the PDF recipe, prefer long options, drop the section rules
Review feedback from Xavier:
- The "grab every PDF" recipe was wrong. `-* +*.pdf` blocks the HTML pages
that carry the PDF links, so the crawl only keeps PDFs linked from the
entry page (verified on a local fixture). There is no PDF-only crawl:
HTTrack finds PDFs by parsing HTML. Reworded to let the site traverse,
with the off-host case handled by a `+host/*.pdf` rule.
- Recipes now use long options throughout. `-%!` in particular is replaced
by `--disable-security-limits`: the bare `!` triggers shell history
expansion and is easy to fumble, and the long name says what it does.
- Dropped the per-section `<hr>` rules; the sibling doc pages don't use
them and the section headings already separate the content.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
* Key the reference tables by long option name, short in parens
Follows the recipe conversion: every table row now leads with the long
option (--depth, --stay-on-same-domain, ...) and carries the short flag in
parentheses. The guru %-flags that have no long form (-%g, -%j, -%o, -%y,
-%z, -%G, -%t, -%N) stay short. Also switches the remaining inline command
snippets in descriptions to long form (--assume, --structure,
--user-agent "").
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
---------
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
--help (and the manual page generated from it) claimed the default socket
count is 8 and the default new-connections-per-second is 10. The engine
actually defaults to 4 sockets (htslib.c) and 5 connections/second, which
is also the -%c ceiling. The 8 in the old text is the -c clamp, not the
default, so the line conflated the two.
Correct the two markers to (*c4) and (*%c5) in htshelp.c and regenerate
man/httrack.1 via man/makeman.sh. html/httrack.man.html carries the same
two-token fix applied directly, since regenerating it needs groff's html
device (Debian: full groff, not groff-base); a later full regen on such a
host reproduces the identical output.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The option parser's `case 'K':` had no `break`, so a bare -K (no trailing
digit) fell through into `case 'c':`, hit its no-digit branch, and forced
the socket count back to the default of 4. `-c8 -K` ended up with 4
connections, not 8; the -K rewrite mode quietly overrode an earlier -c.
Only a trailing bare -K triggered it (`-K -c8` and `-Kc8` were fine),
which is why it went unnoticed.
Add the missing break. The regression test drives the observable that
maxsoc has: a socket count above 8 trips the "limited to 8" security
warning in hts-log.txt, so `-c16` warns and `-c16 -K` warns only if the
16 survives. It fails on the pre-fix build and passes after.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The command-line documentation (cmddoc.html, options.html, and the generated httrack.man.html) was unreachable from the documentation index; they only linked each other. Link cmddoc.html into the "How to Use" list so the whole set is reachable, and replace options.html's stale ~2007 option list with a pointer to the generated man page.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
Adds html/android.html, a step-by-step guide for the HTTrack Android app (install, project, address, options, run, browse, storage), with screenshots of the shipped app and an options section covering all eleven tabs and the settings that are missing or fixed on Android. Linked from the documentation index between the WinHTTrack/WebHTTrack and Fred Cohen entries.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
Bumps the copyright year in the html/ documentation footers (and the inline notice in contact.html) from 2007 to 1998-2026. Footer text only, no content changes; fcguide.html (Fred Cohen, upstream) is left untouched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
Four factual bugs in the shipped html/ help pages: the FAQ's stale "SOCKS? Not yet!" answer (SOCKS5 and HTTP CONNECT proxies have shipped), cache.html's false ">4GiB ZIP64 not supported" claim, two -mime:video/* example rows in filters.html mislabeled "application/", and a compile-breaking fprintf(stder, ...) in plug.html's sample module.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
src/vcpkg.json had no builtin-baseline, so vcpkg resolved OpenSSL/zlib/brotli/zstd in Classic mode from whatever ports tree the build runner happened to sit at: the crypto shipped next to libhttrack.dll was a property of the runner image, not of any commit. Pin the baseline (resolves OpenSSL 3.6.3, zlib 1.3.2#1, what the Windows build already ships) so it becomes deterministic, and add a Dependabot vcpkg entry so the baseline advances on a schedule instead of freezing into a stale, CVE-bearing OpenSSL. The Windows workflow now fetches the pinned baseline before building, read back from the manifest so Dependabot bumps need no workflow edit.
Closes#642
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Fold repo-specific conventions from private notes into the tracked
guide: a C conventions section (htssafe.h *t allocators, HTSEXT_API as
the exported ABI surface, Windows-breakable/POSIX-stable ABI split),
the concrete Latin-1 file list behind the byte-safe-edits rule, and
three test gotchas (register NN_*.test in TESTS, installed-binary PATH
shadowing, set -e in new .test scripts).
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
#581 reports a Windows-only data loss: an interrupted mirror that resumes into an unusable Content-Range 206 drops its partial and never refetches, where the same sequence recovers the file whole on Linux. The trigger is the damaged cache a hard TerminateProcess leaves behind (MSYS can't signal a native exe), which pass 2 has to repair before resuming.
This does not fix the engine. I couldn't reproduce the failure on Linux, so an engine change would be guesswork. What it adds is a deterministic test that drives the damaged-cache regime, plus an analysis of where the two platforms part.
Test 71 leaves a partial and a temp-ref in pass 1, truncates `new.zip` past its last local entry so the central directory a hard kill never wrote is gone and the repair path runs, then resumes into the hostile 206. On Linux that fires the repair, takes the "unusable range -> restart whole" branch, and recovers the file whole every time, across every damage severity I tried, including a repair that recovers zero entries or fails outright. So the cache repair and the restart-whole logic are not themselves where Linux and Windows differ.
Root cause, as far as I can pin it from Linux: restart-whole doesn't refetch. It removes the partial and the temp-ref, flags `STATUSCODE_NON_FATAL`, and leans on the ordinary retry to requeue the URL. The requeued attempt rebuilds the request from the cache, then the temp-ref, then the on-disk partial. On Linux the removals leave none of those, so the retry is a clean whole-file GET and it succeeds. For the file to be lost, the retried attempt has to send a Range again: a second unusable 206, a second -5, and once the retry budget is spent the partial is already gone. That second Range can only come from a temp-ref or partial that outlived the restart-whole removal, which points at a Windows-specific removal or path effect I can't confirm from here.
For the maintainer: the fragile hinge is that restart-whole depends on a budget-consuming retry that re-derives its Range state from disk. A sturdier fix would make the retried attempt refuse to resume, via a per-link "refetch whole, no Range" flag the request builder honors, so a leftover temp-ref or partial can't re-enter the 206 loop whatever the removal quirk turns out to be. Checking that `UNLINK` and `url_savename_refname_remove` actually succeed on Windows would confirm the mechanism first.
Test 71 carries the same Windows skip as test 48. Lifting that skip should reproduce the failure, and turn the test into the fix's verification.
Refs #581.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
webhttrack splits the raw HTTP POST body into argv and calls hts_main2 directly, but that body is in the web form's declared charset (LANGUAGE_CHARSET: ISO-8859-1, windows-125x, BIG5, gb2312, shift-jis, depending on the language), not UTF-8. The httrack CLI and WinHTTrack both hand the engine UTF-8, which htsname's path budget and htscache's format detection now assume, so a non-ASCII output path or URL from the web UI reached the engine as raw form-charset bytes and put the mirror in a mojibake directory instead of the one the user named.
The command line is now converted from the current LANGUAGE_CHARSET to UTF-8 before the crawl starts, in htsserver.c right before webhttrack_main() where that charset is known via LANGSEL. ASCII and already-UTF-8 input pass through untouched. webhttrack's own argv also gets the Windows hts_argv_utf8 treatment the CLI already has.
This exports hts_convertStringToUTF8 so htsserver (which links the shared library) can reach it: an additive ABI change, new symbol, soname unchanged, alongside the already-exported hts_convertStringSystemToUTF8. Flagging it since it touches the public export set.
Test 68 drives the real htsserver over HTTP, posts a start command whose -O dir is café in ISO-8859-1, and checks the mirror lands under the UTF-8 café directory, not the ISO-8859-1 twin. It fails on master and passes with the fix.
Closes#629
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
unzRepair parses the local file header of a damaged cache zip during
repair. READ_32 combined two int-typed READ_16 halves as
READ_16(adr) | (READ_16((adr)+2) << 16); when the high half has bit 15
set (>= 0x8000), shifting it left by 16 exceeds INT_MAX, which is signed
overflow. UBSan aborts on the CRC/size fields of a header whose high
16-bit word has that bit set. Repair runs on hostile input (a corrupt or
foreign new.zip). Cast the halves to uLong before the shift, matching the
sibling minizip readers in unzip.c and zip.c.
Closes#639
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A single `-O` sets both `path_html` and `path_log`, so `httrack -O café url` sends the logs through `path_log` too. argv is UTF-8, so `path_log` holds UTF-8 bytes, but the two-file log branch in `hts_main_internal()` still created its directory and opened `hts-log.txt`/`hts-err.txt` through the ANSI `structcheck()`/`fopen()`. On Windows those read the bytes as the codepage and dropped the logs into a second, mangled `caf<mojibake>/` directory beside the mirror. This routes that branch through the same UTF-8 wrappers #628 used for the mirror root (`structcheck_utf8`/`FOPEN`/`UNLINK`/`fexist_utf8`), so the logs land under `café/` with the mirror. On a UTF-8 filesystem the wrappers resolve to the same calls, so only Windows changes.
This is the log half of #630. I left the cache half out on purpose: relocating `hts-cache` is a much larger, coupled change, because the cache is written through minizip's own ANSI `fopen` and cleaned up through an ANSI `opendir`, so half-converting it would split or break the cache rather than move it. That wants its own PR, and #630 should stay open for it.
Test 69 mirrors into a single non-ASCII `-O` and asserts the audits read `hts-log.txt` from that directory. Like test 64 it only bites on the Windows CI leg, where the two encodings differ; on Linux the change is a no-op and the whole suite stays green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
HTTrack gives a file a temporary `.delayed` placeholder name while it can't resolve the file type yet, then renames the placeholder once the type is known. The problem is that url_savename appends that marker before it enforces the 236-char path ceiling, and the ceiling cuts the tail of the last path segment, which is exactly where the marker lives. Once the trailing `.delayed` is gone the name no longer matches `IS_DELAYED_EXT`, so back_delayed_rename bails out ("nothing bound to the placeholder name") and the downloaded file is never moved to its final name. It is reachable through #133-style deep paths whose final segment is a long hashed filename.
The fix reserves the trailing `.<id>.delayed` in the last-segment copy loop and trims the head of the segment instead, so the result still fits under the ceiling. The hex-id scan stops at its dot separator, so a wholly-hex hashed base is never mistaken for the collision tag and pulled into the marker.
Test 67 drives the naming path over a deep, over-long delayed URL and checks that the marker survives the cut. It fails on master and passes with the fix.
Closes#623
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The test matched engine output with `echo "$out" | grep -q PATTERN`. Under
`set -o pipefail`, grep -q exits the moment it matches, so echo takes a SIGPIPE
writing the rest and the pipeline reports failure even though the pattern was
found. The `|| { echo FAIL; exit 1; }` guard then fired on a passing case,
turning it into an intermittent, output-size-dependent failure (seen on the
Debian buildd leg of #632).
Match with here-strings instead, dropping the pipe and the race entirely.
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The link-rewrite step that drops a redundant :80 read the port with a
hand-rolled digit accumulator, then skipped a hardcoded 3 chars (":80")
when the value equaled 80. Any longer spelling that still evaluates to
80 lost only its first 3 chars and glued the rest onto the host:
http://127.0.0.1:080/x became 127.0.0.10/x, :0080 became 127.0.0.180.
The accumulator was also unchecked, so :4294967376 wrapped to 80 and
took the same path (#614 shape).
Extract the strip into hts_strip_default_port() and parse the matched
digits with the range-checked hts_parse_url_port(): a wrapped or
out-of-range value no longer aliases 80, and a genuine default is
dropped by its full matched length instead of 3. Covered by the
"stripport" engine self-test (tests/66_engine-port80-strip.test).
Closes#627
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-18 07:54:27 +02:00
552 changed files with 40463 additions and 10221 deletions
AC_MSG_ERROR([BASH_SHELL must not contain shell or make metacharacters, got: $BASH_SHELL]) ;;
'' | [[\\/]]* | ?:[[\\/]]*) ;;
*) AC_MSG_ERROR([BASH_SHELL must be an absolute path, got: $BASH_SHELL]) ;;
esac
hts_bash_override=$BASH_SHELL
AC_PATH_PROGS([BASH_SHELL], [bash], [/bin/bash])
# An absolute override is taken verbatim, so BASH_SHELL=/bin/sh would put #895 back and only
# surface at "make check" or "make deb" (#908). What we found ourselves is only a warning:
# a box with no bash still builds, it just cannot run those two.
AC_MSG_CHECKING([whether $BASH_SHELL is a bash outside POSIX mode])
hts_bash_why=
hts_bash_env=no
# AS_EXECUTABLE_P, not "test -x": the PATH search above already demands a regular file, and
# bash blocks forever reading a FIFO it failed to exec, so -x alone hangs configure (#922).
if ! AS_EXECUTABLE_P(["$BASH_SHELL"]); then
hts_bash_why="not an executable regular file"
elif test -z "$("$BASH_SHELL" -c 'echo "${BASH_VERSINFO[[0]]}"' 2>/dev/null)"; then
# Not BASH_VERSION: that is an ordinary variable, so any shell echoes back a spoofed one.
hts_bash_why="not a bash: it reports no BASH_VERSINFO"
else
# sh-mode bash reports a version too, so only SHELLOPTS tells the two apart.
case $("$BASH_SHELL" -c 'echo ":$SHELLOPTS:"' 2>/dev/null) in
*:posix:*)
hts_bash_why="a bash in POSIX sh-mode"
# POSIXLY_CORRECT and an exported SHELLOPTS do that to every bash on the box, so no path
# can pass and blaming this one would send the user hunting for another. Reading them
# here would not do: configure puts its own shell in posix mode, which sets both.
case $(env -u POSIXLY_CORRECT -u SHELLOPTS "$BASH_SHELL" -c 'echo ":$SHELLOPTS:"' 2>/dev/null) in
'' | *:posix:*) ;; # no "env -u", or posix whatever the environment: blame the path
*) hts_bash_env=yes ;;
esac
;;
esac
fi
if test -z "$hts_bash_why"; then
AC_MSG_RESULT([yes])
else
AC_MSG_RESULT([no])
hts_bash_msg="POSIXLY_CORRECT or SHELLOPTS forces every bash into POSIX sh-mode, $BASH_SHELL included. Clear them for configure and for make, which hands them to make check and make deb: env -u POSIXLY_CORRECT -u SHELLOPTS ..."
if test "$hts_bash_env" = yes; then
if test -n "$hts_bash_override"; then
AC_MSG_ERROR([$hts_bash_msg])
fi
AC_MSG_WARN([$hts_bash_msg])
else
if test -n "$hts_bash_override"; then
AC_MSG_ERROR([BASH_SHELL=$BASH_SHELL is $hts_bash_why])
fi
AC_MSG_WARN([no usable bash found: $BASH_SHELL is $hts_bash_why. "make check" and "make deb" need one; pass BASH_SHELL=/path/to/bash])
fi
fi
AC_PROG_CC
AM_PROG_CC_C_O
m4_warn([obsolete],
@@ -49,14 +113,13 @@ m4_warn([obsolete],
# script's behavior did not change. They are probably safe to remove.
AC_CHECK_INCLUDES_DEFAULT
AC_PROG_EGREP
# $(SED) substitutes $(datadir) into src/webhttrack
AC_PROG_SED
LT_INIT
AC_PROG_LN_S
LT_INIT
# bash, used to run the test scripts (see tests/Makefile.am TEST_LOG_COMPILER)
Run one target by hand: `fuzz/fuzz-url -max_total_time=300 corpusdir fuzz/corpus/url`. Seed corpora live in `corpus/<target>/`; a crash reproducer is replayed with `fuzz/fuzz-url crash-file`.
`fuzz-arc` is the odd one out: it drives proxytrack's `.arc` reader the way `--convert` does, through a temp file rather than a buffer, and it compiles `src/proxy/store.c` into the harness because proxytrack does not link libhttrack. Both readers and the writer print to stderr on malformed input, so pass `-close_fd_mask=2` for anything longer than a corpus replay.
This file lists all changes and fixes that have been made for HTTrack
3.49-16
+ New: macOS ships a signed and notarized HTTrack.app with its own icon, bundling the OpenSSL it needs so a downloaded copy launches (#890, #900, #901, #950)
+ New: the desktop icons install into the hicolor theme, scalable SVG included, so Icon=httrack resolves in a launcher (#932, #933)
+ Fixed: an unauthenticated PROPFIND overflowed the ProxyTrack DAV item buffer (#836)
+ Fixed: the cached-headers block was built with an unbounded sprintf (#841)
+ Fixed: a chunked response carrying trailers was rejected as "Invalid chunk" (#855)
+ Fixed: an over-long URL aborted the whole mirror inside the cache instead of reading as a miss (#935, #936)
+ Fixed: ProxyTrack could not re-read the .arc it writes, and crashed on a record whose body it could not read (#834, #929, #931)
+ Fixed: ProxyTrack logged a hashtable stats line and every PROPFIND body to stderr (#911, #918)
+ Fixed: the frozen-slot spool was written inside the mirror namespace, so it landed in the mirrored tree (#859)
+ Fixed: a stack overflow killed httrack with no diagnostic, crash reports named no frame of the executable, and on armhf they had no frames at all (#866, #889, #892)
+ Fixed: an install moved away from its configured prefix could not find its data directory, its shared library, or the html symlink (#885, #887, #894, #906)
+ Fixed: tooltips in the web GUI broke on a translation containing an apostrophe; the escaping now also covers quotes, backslashes, markup and DBCS lead bytes (#864)
+ Fixed: nine strings of the Windows GUI's option dialogs were untranslated in 26 of the 30 language files, and six language files disagreed with their own declared charset (#863, #963)
+ Fixed: the AppStream metainfo still advertised WebHTTrack 3.49.8 (#884)
+ Fixed: the installed development headers did not compile standalone (#943)
+ Fixed: configure discarded a user-supplied BASH_SHELL, resolved bash to /bin/sh on macOS, and hung on one pointing at a FIFO (#891, #895, #908, #922)
+ Changed: the masthead wordmark and the rings background are SVG, so they stay sharp on a hi-DPI screen (#910, #916)
+ Changed: multiple internal hardening, build, test and CI improvements
3.49-15
+ New: --single-file rewrites each saved page with its assets inlined as data: URIs (#713)
+ New: --changes reports what a crawl added, updated or removed against the previous mirror (#714)
+ New: --sitemap and --sitemap-url ingest sitemaps, so pages nothing links to are still found (#712)
+ New: webhttrack exposes --warc-cdx, --wacz and --warc-max-size (#862)
+ Fixed: --update and --purge-old destroyed a good local copy when the re-fetch got no response, was aborted mid-read, or when the backup meant to protect it failed (#746, #748, #758, #775)
+ Fixed: an FTP re-fetch truncated the mirrored file, resumed a complete mirror with REST and spliced the old body into the new one, and a successful transfer was blanked when its backlog slot was swapped out (#771, #797, #798, #823)
+ Fixed: a chunked response cut at a chunk boundary was stored and cached as complete (#840)
+ Fixed: cache repair deleted the old cache before a rename it never checked, and the unlink-then-rename fallback could lose the destination (#779, #786, #790, #824)
+ Fixed: several backward scans from strlen(s) - 1 read before their buffer on an empty string (#730, #768, #770, #814, #821)
+ Fixed: a URL could be saved onto the engine's own temporary files, and the final path segment was never clamped (#774, #842, #852)
+ Fixed: a query-string character reference the page charset cannot represent was left unescaped, changing how the query parses (#854)
+ Fixed: a second --update pass overwrote the previous WARC and regenerated a page-less WACZ (#759)
+ Fixed: WARC output dropped URLs of 1005 bytes or more, archived nothing for a 304 revisit, and marked an engine-forced not-modified as one the server sent (#778, #785, #826, #838, #839)
+ Fixed: ProxyTrack overflowed its .arc header block, walked past the end of an .ndx buffer, trusted an unparsed offset, and crashed on a PROPFIND or on an entry with no usable Last-Modified (#793, #820, #825, #828)
+ Fixed: webhttrack leaked the session id into crawled pages, overflowed its command line so a quoted value could inject flags, let a posted project path repoint the served root, and built its redirect Location in a 256-byte stack buffer (#700, #706, #707, #710)
+ Fixed: in the web GUI, options ticked on by default could not be un-ticked, "max site size" set a per-file cap instead of the HTML one, and the mirror link on the finished page could not be followed (#708, #709, #725)
+ Fixed: htsserver labelled every PNG as image/gif and offered JPEG as a download, and an unauthenticated GET of a directory spun the server forever (#724, #875)
+ Fixed: webhttrack hung at exit when no mirror had been launched (#753)
+ Fixed: oversized cache and header fields aborted the engine or overflowed a neighbouring field instead of being clipped (#701, #715, #717, #722, #732)
+ Fixed: a -%S rules file of 4 GB or more overran the heap (#702)
+ Fixed: the CLI display did not repaint when the terminal was resized (#97)
+ Fixed: a document with no declared charset double-encoded the title lifted for the local index (#848)
+ Fixed: a fragment on an inlined reference was dropped, losing an SVG sprite selector (#766)
+ Fixed: several time helpers handed out libc's shared gmtime/localtime static rather than a reentrant breakdown (#794, #806)
+ Changed: fatal-signal backtraces name engine frames instead of a bare module and offset (#705)
+ Changed: --without-zlib is rejected at configure time rather than failing at link (#735)
+ Changed: the offline documentation gains one GUI guide with screenshots, an Android option reference, and a restructured index
+ Changed: multiple internal hardening, test and CI improvements
3.49-14
+ New: WARC/1.1 archive output (--warc), with a sorted CDXJ index (--warc-cdx) and WACZ packaging (--wacz), also available from webhttrack (#668)
+ New: -%F takes named footer fields such as {url}, {lastmodified}, {mime}, {charset} and {status} instead of a fixed layout (#667)
+ New: a command-line guide organized by task ships with the offline documentation (#649)
+ Fixed: on Windows, paths beyond MAX_PATH truncated files, and mirroring into a long or non-ASCII directory silently failed (#133)
+ Fixed: the top index showed mojibake for non-ASCII project names and categories on Windows (#216)
+ Fixed: webhttrack handed the engine the web form's charset rather than UTF-8, so a non-ASCII path mirrored into a mojibake directory (#629)
+ Fixed: a non-ASCII single -O left the logs and the cache in a mangled twin directory on Windows (#630)
+ Fixed: a path-ceiling truncation dropped the .delayed marker, losing the file (#623)
+ Fixed: an oversized -%F footer aborted the crawl instead of being skipped (#669)
+ Fixed: a rejected 206 resume could loop and lose the file rather than refetch it whole (#581)
+ Fixed: default-port stripping was scheme-blind and dropped explicit ports from https and ftp URLs, and a :80 written with leading zeros mangled the host (#627, #638)
+ Fixed: -K silently reset the -c socket count (#650)
+ Fixed: signed-shift undefined behaviour in the zip-repair local-header read (#639)
+ Changed: the offline documentation drops stale facts, gains an Android help page, and documents the filter wildcards and the real long option forms
+ Changed: multiple internal hardening, test and CI improvements
3.49-13
+ New: SOCKS5 proxy support, with scheme-aware -P URLs (socks5://, socks5h://, connect://) and plain HTTP tunneled through a CONNECT-only proxy (#563, #564)
+ New: decode brotli and zstd content codings, advertised over TLS only as browsers do (#556)
@@ -81,7 +148,7 @@ This file lists all changes and fixes that have been made for HTTrack
+ Fixed: report why a -%L URL list could not be loaded (#49)
+ Changed: multiple internal hardening, build and CI improvements
.49-9
3.49-9
+ Fixed: file-type detection from the Content-Type header: trust a declared type over a binary URL extension, honor --assume under the delayed type check, and keep a known extension against a bogus or empty Content-Type (#267, #29, #56)
+ Fixed: an uninitialized-buffer read when the Content-Type is empty (#411)
+ Fixed: restored C++ source-compatibility of the installed headers so reverse dependencies (httraqt) build again (#413)
<metaname="description"content="HTTrack is an easy-to-use website mirror utility. It allows you to download a World Wide website from the Internet to a local directory,building recursively all structures, getting html, images, and other files from the server to your computer. Links are rebuiltrelatively so that you can freely browse to the local site (works with any browser). You can mirror several sites together so that you can jump from one toanother. You can, also, update an existing mirror site, or resume an interrupted download. The robot is fully configurable, with an integrated help"/>
<metaname="keywords"content="httrack, HTTRACK, HTTrack, winhttrack, WINHTTRACK, WinHTTrack, offline browser, web mirror utility, aspirateur web, surf offline, web capture, www mirror utility, browse offline, local site builder, website mirroring, aspirateur www, internet grabber, capture de site web, internet tool, hors connexion, unix, dos, windows 95, windows 98, solaris, ibm580, AIX 4.0, HTS, HTGet, web aspirator, web aspirateur, libre, GPL, GNU, free software"/>
syslog("user rejected due to too many copy attemps : ".$addr);
}
<pre>
</tt>
<br>
@@ -409,15 +372,15 @@ Example:<br>
FOS('mycompany.com','smith?subject=Hi, John','Click here to email me!')<br>
// --><br>
</script><br>
<noscript><br>
smith at mycompany dot com<br>
</noscript><br>
<noscript><br>
smith at mycompany dot com<br>
</noscript><br>
</tt>
<br>
</li><li>Another one is to create images of emails<br>
Good: Efficient, does not require javascript<br>
Bad: There is still the problem of the link (mailto:), images are bigger than text, and it can cause problems for blind people (a good solution is use an ALT attribute with the email written like "smith at mycompany dot com")<br>
How to do: Not so obvious of you do not want to create images by yourself<br>
How to do: Not so obvious if you do not want to create images by yourself<br>
Example: (php, Unix)<br>
<tt>
@@ -491,7 +454,7 @@ echo <br>
</li><li>You can also create temporary email aliases, each week, for all users<br>
Good: Efficient, and you can give your real email in your reply-to address<br>
Bad: Temporary emails<br>
How to do: Not so hard todo<br>
How to do: Not so hard todo<br>
Example: (script & php, Unix)<br>
<tt>
@@ -566,26 +529,14 @@ And then, put the email address in your pages through:
<metaname="description"content="HTTrack is an easy-to-use website mirror utility. It allows you to download a World Wide website from the Internet to a local directory,building recursively all structures, getting html, images, and other files from the server to your computer. Links are rebuiltrelatively so that you can freely browse to the local site (works with any browser). You can mirror several sites together so that you can jump from one toanother. You can, also, update an existing mirror site, or resume an interrupted download. The robot is fully configurable, with an integrated help"/>
<metaname="keywords"content="httrack, HTTRACK, HTTrack, winhttrack, WINHTTRACK, WinHTTrack, offline browser, web mirror utility, aspirateur web, surf offline, web capture, www mirror utility, browse offline, local site builder, website mirroring, aspirateur www, internet grabber, capture de site web, internet tool, hors connexion, unix, dos, windows 95, windows 98, solaris, ibm580, AIX 4.0, HTS, HTGet, web aspirator, web aspirateur, libre, GPL, GNU, free software"/>
<metaname="description"content="HTTrack is an easy-to-use website mirror utility. It allows you to download a World Wide website from the Internet to a local directory,building recursively all structures, getting html, images, and other files from the server to your computer. Links are rebuiltrelatively so that you can freely browse to the local site (works with any browser). You can mirror several sites together so that you can jump from one toanother. You can, also, update an existing mirror site, or resume an interrupted download. The robot is fully configurable, with an integrated help"/>
<metaname="keywords"content="httrack, HTTRACK, HTTrack, winhttrack, WINHTTRACK, WinHTTrack, offline browser, web mirror utility, aspirateur web, surf offline, web capture, www mirror utility, browse offline, local site builder, website mirroring, aspirateur www, internet grabber, capture de site web, internet tool, hors connexion, unix, dos, windows 95, windows 98, solaris, ibm580, AIX 4.0, HTS, HTGet, web aspirator, web aspirateur, libre, GPL, GNU, free software"/>
<title>HTTrack Website Copier - Cache format specification</title>
For updating purpose, HTTrack stores original (untouched) HTML data,
references to downloaded files, and other meta-data (especially parts of the HTTP headers) in a cache,
located in the hts-cache directory. Because local html pages are always modified to "fit" the local
filesystem structure, and because meta-data such as the last-Modified date and Etag can not be stored
with the associated files, the cache is absolutely mandatory for reprocessing (update/continue) phases.
<br/><br/>
<h3>The (new) cache.zip format</h3>
The 3.31 release of HTTrack introduces a new cache format, more extensible and efficient than the previous one (ndx/dat format).
The main advantages of this cache are:
<ul>
<li>One single file for a complete website cache archive</li>
<li>Standard <ahref="http://www.pkware.com/products/enterprise/white_papers/appnote.txt"target="_new">ZIP</a> format, that can be easily reused on most platforms and languages</li>
<li>Compressed data with the efficient and opened <ahref="http://www.gzip.org/zlib/"target="_new">zlib</a> format</li>
</ul>
The cache is made of ZIP files entries ; with one ZIP file entry per fetched URL (successfully or not - errors are also stored).<br/>
For each entry:
<ul>
<li>The ZIP file name is the original URL [<small><ahref="#orig">see notes below</a></small>]</li>
<li>The ZIP file contents, <b>if available</b>, is the original (compressed, using the deflate algorythm) data</li>
<li>The ZIP file extra field (in the local file header) contains a list of meta-fields, very similar to the <ahref="http://www.ietf.org/rfc/rfc2616.txt?number=2616"target="new_">HTTP</a> headers fields. See also <ahref="http://www.ietf.org/rfc/rfc2396.txt?number=2396"target="new_">RFC</a>.</li><br/>
<li>The ZIP file timestamp follows the "Last-Modified-Since" field given for this URL, if any</li>
There are also specific issues regarding this format:
<ul>
<li>The data in the central directory (such as CD extra field, and CD comments) are not used</li>
<li>The ZIP archive is allowed to contains more than 2^16 files (65535) ; in such case the total number of entries in the 32-bit central directory is 65536 (0xffff), but the presence of the 64-bit central directory is not mandatory</li>
<li>The ZIP archive is allowed to contains more than 2^32 bytes (4GiB) ; in such case the 64-bit central directory must be present <b>(not currently supported)</b></li>
</ul>
<br/>
<b>Meta-data stored in the "extra field" of the local file headers</b><br/>
The extra field is composed of text data, and this text data is composed of distinct lines of headers.
The end of text, <b>or</b> a double CR/LF, mark the end of this zone.
This method allows you to optionally store original HTTP headers just after the "meta-data" headers for informational use.<br/>
<br/>
<b>The status line (the first headers line)</b><br/>
Indicates if the data are present (value=1) in the cache (that is, as ZIP data), or in an external file (value=0).
This field MUST be the first field.
<li>X-StatusCode</li><br>
The modified (by httrack) status code after processing. 304 error codes ("Not modified"), for example, are transformed into "200" codes after processing.
<li>X-StatusMessage</li><br>
The modified (by httrack) status message.
<li>X-Size</li><br>
The stored (either in cache, or in an external file) data size.
<li>X-Charset</li><br>
The original charset.
<li>X-Addr</li><br>
The original URL address part.
<li>X-Fil</li><br>
The original URL path part.
<li>X-Save</li><br>
The local filename, depending on user's "build structure" preferences.
</ul>
<br/>
<b>Standard (RFC 2616) "useful" fields:</b><br/>
<ul>
<li>Content-Type</li>
<li>Last-Modified</li>
<li>Etag</li>
<li>Location</li>
<li>Content-Disposition</li>
</ul>
<br/>
<b>Specific fields in "BNF-like" grammar:</b><br/>
The 3.31 release of HTTrack introduces a new cache format, more extensible and efficient than the previous one (ndx/dat format).
The main advantages of this cache are:
<ul>
<li>One single file for a complete website cache archive</li>
<li>Standard <ahref="http://www.pkware.com/products/enterprise/white_papers/appnote.txt"target="_new">ZIP</a> format, that can be easily reused on most platforms and languages</li>
<li>Compressed data with the efficient and opened <ahref="http://www.gzip.org/zlib/"target="_new">zlib</a> format</li>
</ul>
The cache is made of ZIP files entries ; with one ZIP file entry per fetched URL (successfully or not - errors are also stored).<br/>
For each entry:
<ul>
<li>The ZIP file name is the original URL [<small><ahref="#orig">see notes below</a></small>]</li>
<li>The ZIP file contents, <b>if available</b>, is the original (compressed, using the deflate algorythm) data</li>
<li>The ZIP file extra field (in the local file header) contains a list of meta-fields, very similar to the <ahref="http://www.ietf.org/rfc/rfc2616.txt?number=2616"target="new_">HTTP</a> headers fields. See also <ahref="http://www.ietf.org/rfc/rfc2396.txt?number=2396"target="new_">RFC</a>.</li><br/>
<li>The ZIP file timestamp follows the "Last-Modified-Since" field given for this URL, if any</li>
There are also specific issues regarding this format:
<ul>
<li>The data in the central directory (such as CD extra field, and CD comments) are not used</li>
<li>The ZIP archive is allowed to contains more than 2^16 files (65535) ; in such case the total number of entries in the 32-bit central directory is 65536 (0xffff), but the presence of the 64-bit central directory is not mandatory</li>
<li>The ZIP archive is allowed to contains more than 2^32 bytes (4GiB) ; in such case the 64-bit central directory is emitted automatically (a single stored entry of 4GiB or more is not supported)</li>
</ul>
<br/>
<b>Meta-data stored in the "extra field" of the local file headers</b><br/>
The extra field is composed of text data, and this text data is composed of distinct lines of headers.
The end of text, <b>or</b> a double CR/LF, mark the end of this zone.
This method allows you to optionally store original HTTP headers just after the "meta-data" headers for informational use.<br/>
<br/>
<b>The status line (the first headers line)</b><br/>
Indicates if the data are present (value=1) in the cache (that is, as ZIP data), or in an external file (value=0).
This field MUST be the first field.
<li>X-StatusCode</li><br>
The modified (by httrack) status code after processing. 304 error codes ("Not modified"), for example, are transformed into "200" codes after processing.
<li>X-StatusMessage</li><br>
The modified (by httrack) status message.
<li>X-Size</li><br>
The stored (either in cache, or in an external file) data size.
<li>X-Charset</li><br>
The original charset.
<li>X-Addr</li><br>
The original URL address part.
<li>X-Fil</li><br>
The original URL path part.
<li>X-Save</li><br>
The local filename, depending on user's "build structure" preferences.
</ul>
<br/>
<b>Standard (RFC 2616) "useful" fields:</b><br/>
<ul>
<li>Content-Type</li>
<li>Last-Modified</li>
<li>Etag</li>
<li>Location</li>
<li>Content-Disposition</li>
</ul>
<br/>
<b>Specific fields in "BNF-like" grammar:</b><br/>
<metaname="description"content="HTTrack is an easy-to-use website mirror utility. It allows you to download a World Wide website from the Internet to a local directory,building recursively all structures, getting html, images, and other files from the server to your computer. Links are rebuiltrelatively so that you can freely browse to the local site (works with any browser). You can mirror several sites together so that you can jump from one toanother. You can, also, update an existing mirror site, or resume an interrupted download. The robot is fully configurable, with an integrated help"/>
<metaname="keywords"content="httrack, HTTRACK, HTTrack, winhttrack, WINHTTRACK, WinHTTrack, offline browser, web mirror utility, aspirateur web, surf offline, web capture, www mirror utility, browse offline, local site builder, website mirroring, aspirateur www, internet grabber, capture de site web, internet tool, hors connexion, unix, dos, windows 95, windows 98, solaris, ibm580, AIX 4.0, HTS, HTGet, web aspirator, web aspirateur, libre, GPL, GNU, free software"/>
<p>With no other options HTTrack mirrors that site, stays on the same host, follows
links to any depth, rebuilds them to browse offline, and stores everything under
<tt>mydir</tt>. The same directory also holds the log files and the
<tt>hts-cache/</tt> folder that makes a later update or resume possible.</p>
<p>Two defaults are worth knowing up front, because both catch people out:</p>
<ul>
<li>HTTrack throttles itself to about <b>100 KB/s</b> even when you pass no rate
option. If a mirror feels slow, that is why. See
<ahref="#limits">Limits</a> for how to lift it.</li>
<li>The download proceeds as a well-behaved robot: it identifies itself as
<tt>HTTrack</tt>, obeys <tt>robots.txt</tt>, and sends a Referer with each
request. A site that blocks that behavior needs the levers in
<ahref="#identity">Identity</a>, not brute force.</li>
</ul>
<h3id="scope">2. Scope: how far the crawl reaches</h3>
<p>Scope decides which links HTTrack is even willing to follow, before any filter
you write. Get this right and most "it downloaded too much" or "it only grabbed
the index" problems disappear.</p>
<tableclass="tblRegular tableWidth"border="0">
<trclass="head"><td><b>Option</b></td><td><b>What it controls</b></td></tr>
<tr><td><tt>--depth (-r)</tt></td><td>Maximum link depth. <b>The start page is level 1</b>, so one level of links out is <tt>-r2</tt>, not <tt>-r1</tt>.</td></tr>
<tr><td><tt>--stay-on-same-address (-a), --stay-on-same-domain (-d), --stay-on-same-tld (-l), --go-everywhere (-e)</tt></td><td>How far off the starting host the crawl may travel: same address (host), same principal domain, same top-level domain (for example .com), or everywhere. The default keeps you on the starting host.</td></tr>
<tr><td><tt>--can-go-down (-D), --can-go-up (-U), --stay-on-same-dir (-S), --can-go-up-and-down (-B)</tt></td><td>Directory travel: down into subdirectories only, up to parent directories only, stay in the same directory, or both up and down.</td></tr>
<tr><td><tt>--near (-n)</tt></td><td>Also fetch non-HTML files "near" a followed link, such as an image linked from a page you kept but hosted elsewhere.</td></tr>
<tr><td><tt>--ext-depth (-%e)</tt></td><td>How many levels of external links to follow once the crawl leaves your scope (default 0).</td></tr>
<tr><td><tt>--test (-t)</tt></td><td>Also HEAD-test links that fall outside the scope, which are normally refused, without downloading them: a way to see what scope is excluding.</td></tr>
<tr><td><tt>--sitemap (-%m), --sitemap-url URL (-%mu)</tt></td><td>Also take start URLs from the site's sitemap, for pages nothing links to. Off by default.</td></tr>
</table>
<p>Link-following only finds what something links to. Anything a site publishes
solely in its sitemap is invisible to HTTrack unless you ask for it.
<tt>--sitemap</tt> reads the start host's <tt>robots.txt</tt> for
<tt>Sitemap:</tt> lines and falls back to <tt>/sitemap.xml</tt>;
<tt>--sitemap-url</tt> names one directly. Nested <tt>sitemapindex</tt> files
and gzipped <tt>.xml.gz</tt> sitemaps are followed. The URLs found become start
URLs with the full depth budget, but they still go through your filters and
scope rules, so a sitemap cannot widen a crawl you deliberately narrowed. It is
off by default because a sitemap can list thousands of pages nothing links
to.</p>
<p>One surprise worth knowing: a sitemap you name with <tt>--sitemap-url</tt>,
and one the site itself declares in <tt>robots.txt</tt>, are fetched even when
<tt>robots.txt</tt> disallows that path, because naming or declaring a sitemap
is an invitation to read it. Only the guessed <tt>/sitemap.xml</tt> obeys a
<tt>Disallow</tt>. The URLs listed inside are gated normally either way.</p>
<p>The single most common surprise is "only the home page came down." That is
usually not a scope option at all: it is an off-host redirect. A start URL of
<tt>http://example.com/</tt> that redirects to <tt>https://www.example.com/</tt>
lands you on a different host, and same-host scope stops the crawl there. Start
from the final URL, or add a filter that re-admits the real host (see
<ahref="#filters">Filters</a>). The log will show the redirect.</p>
<p><tt>-n</tt> is the fix for pages that render locally without their images or
stylesheets: it lets HTTrack pull in requisites that sit just outside scope. Note
that its embedded-asset handling (following <tt>img</tt>, <tt>link</tt>,
<tt>script</tt>, <tt>style</tt> and HTML5 <tt>source</tt>/<tt>track</tt> targets
past the normal depth and filter limits) applies only when <tt>-n</tt> is on; it
is not automatic. It can also over-fetch by dragging in a whole external host from
a single link, in which case name the assets you want with a filter instead.</p>
<h3id="filters">3. Filters and scan rules</h3>
<p>Filters are the number-one source of confusion, and also the tool that solves
most scope problems once you understand them. A filter is a rule that accepts
(<tt>+</tt>) or rejects (<tt>-</tt>) URLs by pattern. The sign is mandatory:
<tt>+pattern</tt> adds, <tt>-pattern</tt> removes, and a bare pattern is an error.</p>
<p>The rules that matter:</p>
<ul>
<li><b>Last match wins.</b> Rules are applied in order and the last one that
matches a URL decides its fate. Order your rules from general to specific.</li>
<li><b>Wildcards.</b><tt>*</tt> matches any run of characters;
<tt>*[a-z]</tt>, <tt>*[0-9]</tt> and similar classes match sets. So
<tt>+*.pdf</tt> means "any URL ending in .pdf".</li>
<li><b>Whitelisting.</b> To keep one site and nothing else, deny everything then
re-admit the host: <tt>"-*" "+example.com/*"</tt>. A lone <tt>+</tt> rule only
adds to the default scope; it never restricts.</li>
<li><b>Size rules.</b><tt>*[>100000]</tt> and <tt>*[<1000]</tt> filter by
byte size. Because size is only known once the transfer starts, an oversize file
is fetched partway and then aborted, not skipped for free.</li>
<li><b>mime: rules.</b> A rule like <tt>-mime:video/*</tt> matches the
<tt>Content-Type</tt>. That type is only known after the response headers arrive,
so a mime rule <b>cannot stop a request</b>; it can only abort the body. Use a
URL pattern when you want to avoid the fetch entirely.</li>
</ul>
<p>Quote your filters. Shells treat <tt>*</tt>, <tt>[</tt> and sometimes <tt>+</tt>
specially, so wrap each rule in quotes as shown above. The full pattern language,
with tables for wildcards, size and mime, is in
<ahref="filters.html">the filters page</a>, and the
<ahref="faq.html">FAQ</a> has a worked tutorial.</p>
<p><b>robots.txt.</b> By default HTTrack obeys <tt>robots.txt</tt> (<tt>-s2</tt>).
<tt>-s0</tt> ignores it entirely, <tt>-s1</tt> obeys it but lets one of your
<tt>+</tt> filters override a disallow for a URL you explicitly asked for. Note
that a <tt>403 Forbidden</tt> is a server refusal, not a robots rule: robots
options will not help there. That is an
<ahref="#identity">identity</a> problem.</p>
<h4>Filter wildcards</h4>
<p>Inside a filter pattern, <tt>*</tt> matches any run of characters; a few
bracket forms match narrower sets. The full table, with size and mime rules, is on
<tr><td><tt>*</tt></td><td>any run of characters</td><td><tt>+*.pdf</tt>— any URL ending <tt>.pdf</tt></td></tr>
<tr><td><tt>*[file]</tt>, <tt>*[name]</tt></td><td>one path segment (any char but <tt>/</tt> and <tt>?</tt>)</td><td><tt>example.com/*[file]/</tt>— a directory-index page</td></tr>
<tr><td><tt>*[path]</tt></td><td>a path, slashes allowed (any char but <tt>?</tt>)</td><td><tt>example.com/*[path].zip</tt></td></tr>
<tr><td><tt>*[param]</tt></td><td>an optional query string</td><td><tt>page.html*[param]</tt> matches with or without <tt>?...</tt></td></tr>
<tr><td><tt>*[a,b,c]</tt></td><td>any one character in the set</td><td><tt>*[a,b,c].txt</tt></td></tr>
<tr><td><tt>*[a-z]</tt></td><td>any one character in the range</td><td><tt>img*[0-9].gif</tt></td></tr>
<tr><td><tt>*[\x]</tt></td><td>the literal character x (escapes <tt>* [ ] \</tt>)</td><td><tt>*[\*]</tt> matches a real <tt>*</tt></td></tr>
<tr><td><tt>*[<NN]</tt>, <tt>*[>NN]</tt></td><td>file size in KB below / above NN</td><td><tt>-*.gif*[<5]</tt> skips GIFs under 5 KB</td></tr>
<tr><td><tt>*[]</tt></td><td>end anchor: nothing may follow</td><td><tt>*.html*[]</tt> rejects <tt>i.html?p=1</tt></td></tr>
</table>
<h3id="limits">4. Limits and politeness</h3>
<p>HTTrack ships cautious on purpose: it is easy to hammer a small site by
accident, and the <ahref="abuse.html">abuse page</a> is worth a read. The limits
below let you go faster when you own the target, and slower when you do not.</p>
<tableclass="tblRegular tableWidth"border="0">
<trclass="head"><td><b>Option</b></td><td><b>What it controls</b></td></tr>
<tr><td><tt>--max-rate (-A)</tt></td><td>Maximum transfer rate in bytes/sec. <b>The default is about 100 KB/s even without this flag.</b> Raise it to go faster.</td></tr>
<tr><td><tt>--sockets (-c)</tt></td><td>Number of parallel connections (default 4). <tt>--tiny</tt>, <tt>--wide</tt> and <tt>--ultrawide</tt> are presets.</td></tr>
<tr><td><tt>--connection-per-second (-%c)</tt></td><td>New connections opened per second (default 5).</td></tr>
<tr><td><tt>--max-size (-M)</tt></td><td>Stop after N bytes <b>received from the network</b> across the whole mirror (this counts what was transferred, not what was saved).</td></tr>
<tr><td><tt>--max-time (-E)</tt></td><td>Stop after N seconds of wall-clock time.</td></tr>
<tr><td><tt>--timeout (-T), --retries (-R), --min-rate (-J), --host-control (-H)</tt></td><td>Idle timeout, retry count, minimum acceptable rate, and host-ban behavior for slow or dead hosts.</td></tr>
<tr><td><tt>--max-pause (-G), --pause (-%G)</tt></td><td>Pause the mirror at N bytes, or pause between files, to spread the load.</td></tr>
</table>
<p><b>The security clamps.</b> To keep an accidental typo from turning into a flood,
HTTrack silently caps a few values: at most 8 connections (<tt>-c</tt>), at most
10 MB/s (<tt>-A</tt>), and at most 5 new connections per second (<tt>-%c</tt>).
Ask for more and you get the ceiling, quietly. The single flag
<tt>--disable-security-limits</tt> lifts all three (the short form <tt>-%!</tt>
also works, but the bare <tt>!</tt> is awkward to type safely in a shell). Use it
only against infrastructure you are allowed to load that hard.</p>
<h3id="names">5. File names and types</h3>
<p>Where local files land, and what they are called, is controlled by the naming
options. This is the second-biggest source of "why did it do that" questions,
usually about a URL like <tt>/article?id=42</tt> or a <tt>.php</tt> page that is
really HTML.</p>
<tableclass="tblRegular tableWidth"border="0">
<trclass="head"><td><b>Option</b></td><td><b>What it controls</b></td></tr>
<tr><td><tt>--structure (-N)</tt></td><td>The local path and name layout. Presets are numeric, and you can also give a template such as <tt>--structure "%h%p/%n%q.%t"</tt>.</td></tr>
<tr><td><tt>--long-names (-L)</tt></td><td>Long names, 8.3 names, or ISO9660 for CD masters.</td></tr>
<tr><td><tt>--assume (-%A)</tt></td><td>Assume a MIME type for an extension, for example <tt>--assume php=text/html</tt>. This also skips the extra HEAD probe HTTrack would otherwise send to learn the type.</td></tr>
<tr><td><tt>--delayed-type-check (-%N), --cached-delayed-type-check (-%D), --check-type (-u), -%t</tt></td><td>When and how the content type is checked, and whether the original extension is kept.</td></tr>
<tr><td><tt>--include-query-string (-%q), --strip-query (-%g)</tt></td><td>Whether the query string appears in the local filename, and whether query keys are stripped when deciding if two URLs are the same file.</td></tr>
</table>
<p>The <tt>-N</tt> presets are built from modular arithmetic on the name fields, so
undocumented number combinations often "work" by accident. If you care about the
exact layout, use an explicit template (the <tt>%h %p %n %q %t</tt> placeholders)
rather than a magic number, and check the result on a small crawl first.</p>
<p>A dynamic page served as <tt>.php</tt> or <tt>.asp</tt> that is actually HTML is
the classic case: without help it can be saved with an extension a browser will
not open locally. <tt>--assume php=text/html</tt> fixes both the extension and the
naming.</p>
<h3id="links">6. Links and page building</h3>
<p>After a page is downloaded, HTTrack parses it for more links and rewrites the
ones it kept so the local copy browses offline. These options tune both halves.</p>
<tableclass="tblRegular tableWidth"border="0">
<trclass="head"><td><b>Option</b></td><td><b>What it controls</b></td></tr>
<tr><td><tt>--keep-links (-K)</tt></td><td>How links are rewritten in saved pages. The numbering is inverted from what you might guess: bare <tt>-K</tt> keeps <b>absolute</b> URLs, and <tt>-K0</tt> is the <b>relative</b> default. <tt>-K3</tt> keeps absolute URIs, <tt>-K4</tt> keeps the original links.</td></tr>
<tr><td><tt>--replace-external (-x), --generate-errors (-o)</tt></td><td>Replace external links with an error page, and generate an error page for links that failed.</td></tr>
<tr><td><tt>--preserve (-%p), --disable-passwords (-%x)</tt></td><td>Leave HTML untouched (no rewriting), and strip passwords out of saved links.</td></tr>
<tr><td><tt>--extended-parsing (-%P), --parse-java (-j)</tt></td><td>Aggressive link discovery, and how much script content is parsed for links.</td></tr>
<tr><td><tt>--mime-html (-%M)</tt></td><td>Save the whole mirror as a single MIME-encapsulated <tt>.mht</tt> archive (<tt>index.mht</tt>).</td></tr>
<tr><td><tt>--single-file (-%Z), --single-file-max-size N</tt></td><td>Once the mirror is finished, rewrite every saved page with its stylesheets, scripts, images and fonts embedded as <tt>data:</tt> URIs. Assets over the cap (10 MB by default) keep their link, as do audio, video, and the links from one page to another. A sibling of <tt>-%M</tt>, not a replacement: see the recipe below for which to pick.</td></tr>
<tr><td><tt>--index (-I), --build-top-index (-%i), --search-index (-%I)</tt></td><td>Build a per-mirror index, a top index across projects, and a searchable keyword index.</td></tr>
</table>
<p>HTTrack finds links by parsing HTML and CSS. It does not run JavaScript, so any
URL a page builds at runtime in script (a lazy-loaded image, a
JavaScript-assembled path) is invisible to the crawler and will be missing from
the mirror. There is no flag that fixes this; the asset has to appear in the
static HTML or CSS to be found. <tt>-%P</tt> widens discovery for links that are
present but awkwardly formatted, not for links that do not exist until script
runs.</p>
<h3id="identity">7. Identity, cookies and login</h3>
<p>By default HTTrack is an honest robot: it sends a <tt>User-Agent</tt> of
<tt>HTTrack</tt>, a Referer with each link (which reveals the crawl path to the
server), and obeys robots. Plenty of sites filter exactly that profile. These
options control what HTTrack says about itself.</p>
<tableclass="tblRegular tableWidth"border="0">
<trclass="head"><td><b>Option</b></td><td><b>What it controls</b></td></tr>
<tr><td><tt>--user-agent (-F)</tt></td><td>The <tt>User-Agent</tt>. Set a browser string to get past crawler blocks; <tt>--user-agent ""</tt> sends none.</td></tr>
<tt>{lastmodified}</tt> (the page's Last-Modified), <tt>{version}</tt>,
<tt>{mime}</tt>, <tt>{charset}</tt>, <tt>{status}</tt> and <tt>{size}</tt>; write
<tt>{{</tt> or <tt>}}</tt> for a literal brace. A footer that contains <tt>%s</tt>
keeps the older positional form (host, path, date in that order). Example:
<tt>-%F "<!-- Mirrored from {url} on {date} -->"</tt>.</p>
<p><b>Login.</b> For HTTP Basic auth, put the credentials in the URL:
<tt>http://user:pass@host/</tt>. An <tt>@</tt> inside the username must be written
<tt>%40</tt>. Only Basic is supported, not Digest.</p>
<p>For cookie or form logins, the simplest path is to log in with a browser, export
its <tt>cookies.txt</tt>, and drop that file in the project directory so HTTrack
sends the session cookie. For a form that needs a POST, <tt>--catchurl</tt> can
capture the exact request your browser sends and replay it. A few cookie caveats
to know: expiry is ignored, there is a silent cap of about 8 cookies sent per
request, and <tt>-b0</tt> disables cookies and the reuse of Basic credentials
across links at the same time.</p>
<h3id="proxy">8. Proxy and network</h3>
<tableclass="tblRegular tableWidth"border="0">
<trclass="head"><td><b>Option</b></td><td><b>What it controls</b></td></tr>
<tr><td><tt>--proxy (-P)</tt></td><td>Route through a proxy. HTTP, SOCKS5 and CONNECT are supported: <tt>-P host:8080</tt>, <tt>-P socks5://host:1080</tt>, <tt>-P connect://host:443</tt>, with optional <tt>user:pass@</tt>.</td></tr>
<tr><td><tt>--httpproxy-ftp (-%f)</tt></td><td>Send FTP requests through the HTTP proxy.</td></tr>
<tr><td><tt>--protocol (-@i)</tt></td><td>Prefer IPv4 or IPv6.</td></tr>
<tr><td><tt>--http-10 (-%h), --keep-alive (-%k), --disable-compression (-%z)</tt></td><td>Force HTTP/1.0 (drops keep-alive and compression, useful for fragile CGI), toggle keep-alive, and toggle compression.</td></tr>
<tr><td><tt>--bind (-%b), --tolerant (-%B)</tt></td><td>Bind to a local address, and accept technically-bogus responses some servers send.</td></tr>
host names at the proxy (remote DNS) for both <tt>socks5://</tt> and
<tt>socks5h://</tt>, so your local resolver is never consulted. And HTTrack does
not verify TLS certificates: HTTPS gives you an encrypted transport, but not an
authenticated one. That is a deliberate choice for a mirroring tool, not a bug,
but it is worth knowing if you are relying on it for trust.</p>
<h3id="update">9. Update and cache</h3>
<p>Every project keeps a cache under <tt>hts-cache/</tt>. It records every URL that
was fetched, together with the options you used, and it is what makes resuming and
updating possible. It is not a size-limited scratch area you can delete: throw it
away and you lose the ability to continue or update the mirror.</p>
<tableclass="tblRegular tableWidth"border="0">
<trclass="head"><td><b>Option</b></td><td><b>What it controls</b></td></tr>
<tr><td><tt>--continue</tt></td><td>Carry on an interrupted mirror, trusting the cache: it does not re-check pages already stored.</td></tr>
<tr><td><tt>--update</tt></td><td>Re-run the mirror, revalidating each page with the server (If-Modified-Since / If-None-Match) and downloading only what changed.</td></tr>
<tr><td><tt>--purge-old=0 (-X0)</tt></td><td>Do not purge. By default an update deletes local files that are no longer part of the mirror; <tt>--purge-old=0</tt> keeps them.</td></tr>
<tr><td><tt>--cache (-C)</tt></td><td>Cache mode. The default already does the right thing and switches to update-checking when it detects an existing mirror.</td></tr>
<tr><td><tt>--debug-cache (-#C), --repair-cache (-#R), --clean</tt></td><td>Inspect the cache, repair its ZIP, and erase cache plus logs.</td></tr>
</table>
<p><b>The purge trap.</b> An <tt>--update</tt> run rebuilds the list of files the
mirror should contain, then deletes any previously-mirrored file that is not on the
new list. This is what keeps a mirror in sync with a shrinking site, but it means a
partial or interrupted update can delete files you meant to keep. If an update might
not complete cleanly, add <tt>-X0</tt> to protect the existing tree, and expect
dynamic pages to look "changed" on every run. See the
<ahref="cache.html">cache page</a> for the details.</p>
<h3id="experts">10. Experts and scripting</h3>
<p>HTTrack is also a scriptable fetch-and-scan tool. These options turn off the
mirror behavior and expose the engine.</p>
<tableclass="tblRegular tableWidth"border="0">
<trclass="head"><td><b>Option</b></td><td><b>What it controls</b></td></tr>
<tr><td><tt>--get URL</tt></td><td>Fetch a single file and stop. Cache, index, depth, cookies and robots are all off for this mode.</td></tr>
<tr><td><tt>--spider --testlinks --skeleton</tt></td><td>Scan without saving, test links at depth 1, or keep HTML only. Handy for checking a site before a real crawl.</td></tr>
<tr><td><tt>--userdef-cmd (-V)</tt></td><td>Run a shell command on each downloaded file; <tt>$0</tt> is the file path. Good for on-the-fly processing.</td></tr>
<tr><td><tt>--callback (-%W)</tt></td><td>Load an external callback module to hook the engine.</td></tr>
<tr><td><tt>--do-not-log (-Q), --quiet (-q), --verbose (-v), --file-log (-f), --extra-log (-z), --debug-log (-Z)</tt></td><td>Logging: quiet, no questions, verbose on screen, and the various log-to-file levels.</td></tr>
</table>
<p>The <ahref="dev.html">developer page</a> covers the callback API and batch use
in more depth.</p>
<h3id="recipes">11. Recipes</h3>
<p>Copy-ready command lines for the tasks people ask about most. Each has the one
<metaname="description"content="HTTrack is an easy-to-use website mirror utility. It allows you to download a World Wide website from the Internet to a local directory,building recursively all structures, getting html, images, and other files from the server to your computer. Links are rebuiltrelatively so that you can freely browse to the local site (works with any browser). You can mirror several sites together so that you can jump from one toanother. You can, also, update an existing mirror site, or resume an interrupted download. The robot is fully configurable, with an integrated help"/>
<metaname="keywords"content="httrack, HTTRACK, HTTrack, winhttrack, WINHTTRACK, WinHTTrack, offline browser, web mirror utility, aspirateur web, surf offline, web capture, www mirror utility, browse offline, local site builder, website mirroring, aspirateur www, internet grabber, capture de site web, internet tool, hors connexion, unix, dos, windows 95, windows 98, solaris, ibm580, AIX 4.0, HTS, HTGet, web aspirator, web aspirateur, libre, GPL, GNU, free software"/>
<metaname="description"content="HTTrack is an easy-to-use website mirror utility. It allows you to download a World Wide website from the Internet to a local directory,building recursively all structures, getting html, images, and other files from the server to your computer. Links are rebuiltrelatively so that you can freely browse to the local site (works with any browser). You can mirror several sites together so that you can jump from one toanother. You can, also, update an existing mirror site, or resume an interrupted download. The robot is fully configurable, with an integrated help"/>
<metaname="keywords"content="httrack, HTTRACK, HTTrack, winhttrack, WINHTTRACK, WinHTTrack, offline browser, web mirror utility, aspirateur web, surf offline, web capture, www mirror utility, browse offline, local site builder, website mirroring, aspirateur www, internet grabber, capture de site web, internet tool, hors connexion, unix, dos, windows 95, windows 98, solaris, ibm580, AIX 4.0, HTS, HTGet, web aspirator, web aspirateur, libre, GPL, GNU, free software"/>
<metaname="description"content="HTTrack is an easy-to-use website mirror utility. It allows you to download a World Wide website from the Internet to a local directory,building recursively all structures, getting html, images, and other files from the server to your computer. Links are rebuiltrelatively so that you can freely browse to the local site (works with any browser). You can mirror several sites together so that you can jump from one toanother. You can, also, update an existing mirror site, or resume an interrupted download. The robot is fully configurable, with an integrated help"/>
<metaname="keywords"content="httrack, HTTRACK, HTTrack, winhttrack, WINHTTRACK, WinHTTrack, offline browser, web mirror utility, aspirateur web, surf offline, web capture, www mirror utility, browse offline, local site builder, website mirroring, aspirateur www, internet grabber, capture de site web, internet tool, hors connexion, unix, dos, windows vista, windows seven, windows 8, solaris, ibm580, AIX 4.0, HTS, HTGet, web aspirator, web aspirateur, libre, GPL, GNU, free software"/>
<li>In case of troubles/problems during transfer, <b><u><fontcolor="red">first check the hts-log.txt (and hts-err.txt) files to figure out what happened</b></u></font>. These log files report all
<li>In case of troubles/problems during transfer, <b><u>first check the hts-log.txt (and hts-err.txt) files to figure out what happened</b></u>. These log files report all
events that may be useful to detect a problem. You can also ajust the debug level of the log files in the option
</li><li>
The tutorial written by Fred Cohen is a very good document to read, to understand how to use the engine,
You have noticed the <tt>-</tt> in the begining of the third rule: this means "refuse links matching the rule"
; and the rule is "any files begining with <tt>www.example.com/gallery/trees/hugetrees/</tt><br>
You have noticed the <tt>-</tt> in the beginning of the third rule: this means "refuse links matching the rule"
; and the rule is "any files beginning with <tt>www.example.com/gallery/trees/hugetrees/</tt><br>
Voila! With these three rules, you have precisely defined what you wanted to capture.<br>
<br>
@@ -361,12 +323,12 @@ You can freely download it, without paying any fees, copy it to your friends, an
There are NO official/authorized resellers, because HTTrack is <b>NOT</b> a commercial product.
But you can be charged for duplication fees, or any other services (example: software CDroms or shareware collections, or fees for maintenance),
but you should have been informed that the software was free software/GPL, and you <b><u>MUST</u></b> have received a copy of the GNU General Public License.
Otherwise this is dishonnest and unfair (ie. selling httrack on ebay without telling that it was a free software is a scam).
Otherwise this is dishonest and unfair (ie. selling httrack on ebay without telling that it was a free software is a scam).
</em>
<br><br><aNAME="QG0b">Q: <strong>Are there any risks of viruses with this software?</strong></a><br>
A: <em>For the software itself:
All official releases (at httrack.com) are checked against all known viruses, and the packaging process is also checked. Archives are stored on Un*x servers, not really concerned by viruses. It has been reported, however, that some rogue freeware sites are embedding free softwares and freewares inside badware installers. Always download httrack from the main site (www.httrack.com), and never from an untrusted source!<br>
All official releases (at httrack.com) are checked against all known viruses, and the packaging process is also checked. Archives are stored on Un*x servers, not really concerned by viruses. It has been reported, however, that some rogue freeware sites are embedding free software and freeware inside badware installers. Always download httrack from the main site (www.httrack.com), and never from an untrusted source!<br>
For files you are downloading on the WWW using HTTrack: You may encounter websites which were corrupted by viruses, and downloading data on these websites might be dangerous if you execute downloaded executables, or if embedded pages contain infected material (as dangerous as if using a regular Browser). Always ensure that websites you are crawling are safe.
(Note: remember that using an antivirus software is a good idea once you are connected to the Internet)</em>
@@ -376,11 +338,8 @@ A: <em>That's right. You can, however, install WinHTTrack on your own machine, a
<br><br><aNAME="QG2">Q: <strong>Where can I find French/other languages documentation?</strong></a><br>
A: <em>Windows interface is available on several languages, but not yet the documentation!</em>
<br><br><aNAME="QG3b">Q: <strong>Is HTTrack working on Windows Vista/Windows Seven/Windows 8 ?</strong></a><br>
A: <em>Yes, it does</em>
<br><br><aNAME="QG3">Q: <strong>Is HTTrack working on Windows 95/98 ?</strong></a><br>
A: <em>No, not anymore. You may try to pick an older release (such as 3.33)</em>
<br><br><aNAME="QG3">Q: <strong>Which systems does HTTrack run on?</strong></a><br>
A: <em>HTTrack runs on current Windows, Linux and other Unix-like systems, and macOS. Very old platforms such as Windows 95/98 are no longer supported; you may try an older release (such as 3.33) on those.</em>
<br><br><aNAME="QG4">Q: <strong>What's the difference between HTTrack, WinHTTrack and WebHTTrack?</strong></a><br>
A: <em>WinHTTrack is the Windows GUI release of HTTrack (with a native graphic shell) and WebHTTrack is the Linux/Posix release of HTTrack (with an html graphic shell)</em>
@@ -394,7 +353,7 @@ A: <em>It should. The <tt>configure.ac</tt> may be modified in some cases, howev
<br><br><aNAME="QG7">Q: <strong>I use HTTrack for professional purpose. What about restrictions/license fee?</strong></a><br>
A: <em>HTTrack is covered by the GNU General Public License (GPL). There is no restrictions using HTTrack for professional purpose,
except if you develop a software which uses HTTrack components (parts of the source, or any other component).
See the <tt>license.txt</tt> file for more information</em>. See also the next question regarding copyright issues when reditributing downloaded material.
See the <tt>license.txt</tt> file for more information</em>. See also the next question regarding copyright issues when redistributing downloaded material.
<br><br><aNAME="QG7b">Q: <strong>Is there any license royalties for distributing a mirror made with HTTrack?</strong></a><br>
A: <em>On the HTTrack side, no. However, sharing, publishing or reusing copyrighted material downloaded from a site requires the authorization of the copyright holders, and possibly paying royalty fees. Always ask the authorization before creating a mirror of a site, even if the site appears to be royalty-free and/or without copyright notice.</em>
@@ -415,7 +374,7 @@ There are several reasons (and solutions) for a mirror to fail. Reading the log
<ul>
<li>Links within the site refers to external links, or links located in another (or upper) directories, not captured by default - the use of filters is generally THE solution, as this is one of the powerful option in HTTrack. <u>See the above questions/answers</u>.</li>
<li>Website <ahref="#Q1b1">'robots.txt' rules</a> forbide access to several website parts - you can disable them, but only with great care!</li>
<li>Website <ahref="#Q1b1">'robots.txt' rules</a> forbid access to several website parts - you can disable them, but only with great care!</li>
<li>HTTrack is filtered (by its default User-agent IDentity) - you can change the Browser User-Agent identity to an anonymous one (MSIE, Netscape..) - here again, use this option with care, as this measure might have been put to avoid some bandwidth abuse (see also the <ahref="abuse.html">abuse faq</a>!)</li>
</ul>
@@ -457,14 +416,14 @@ A: <em>Yes, HTTrack does support (since 3.20 release) ipv6 sites, using A/AAAA e
A: <em>Check the build options (you may have selected user-defined structure with wrong parameters!)</em>
<br><br><aNAME="QT5">Q: <strong>When capturing real audio/video links (.ram), I only get a shortcut!</a></strong></a></br>
A: <em>Yes, but .ra/.rm associated file should be captured together - except if rtsp:// protocol is used (not supported by HTTrack yet), or if proper filters are needed</em>
A: <em>The .ra/.rm associated file can be captured together with the shortcut, if proper filters are set. Streaming protocols such as rtsp:// are out of scope: HTTrack is an HTTP/FTP mirror and does not capture rtsp streams.</em>
<br><br><aNAME="QT6">Q: <strong>Using user:password@address is not working!</a></strong></a></br>
A: <em>Again, first check the <tt>hts-log.txt</tt> and <tt>hts-err.txt</tt> error log files - this can give you precious information<br>
The site may have a different authentication scheme - form based authentication, for example.
In this case, use the URL capture features of HTTrack, it might work.
<br>Note: If your username and/or password contains a '<tt>@</tt>' character, you may have to replace all '<tt>@</tt>'
occurences by '<tt>%40</tt>' so that it can work, such as in <tt>user%40domain.com:foobar@www.foo.com/auth/.
occurrences by '<tt>%40</tt>' so that it can work, such as in <tt>user%40domain.com:foobar@www.foo.com/auth/.
You may have to do the same for all "special" characters like spaces (%20), quotes (%22)..</tt>
</em>
<br><br>
@@ -520,7 +479,7 @@ These rules, stored in a file called robots.txt, are given by the website, to sp
- for example, /cgi-bin or large images files.
They are followed by default by HTTrack, as it is advised. Therefore, you may miss some files that would have been downloaded without
these rules - check in your logs if it is the case:<br>
<tt>Info: Note: due to www.foobar.com remote robots.txt rules, links begining with these path will be forbidden: /cgi-bin/,/images/ (see in the options to disable this)
<tt>Info: Note: due to www.foobar.com remote robots.txt rules, links beginning with these path will be forbidden: /cgi-bin/,/images/ (see in the options to disable this)
</tt>
<br>
If you want to disable them, just change the corresponding option in the option list! (but only disable this option with great care,
@@ -540,7 +499,7 @@ HTTrack must find one. Therefore, two index.html will be produced, one with the
<br>
It might be a good idea to consider that http://www.foobar.com/ and http://www.foobar.com/index.html are the same links, to avoid
duplicate files, isn't it?
NO, because the top index (/) can refer to ANY filename, and if index.html is generally the default name, index.htm can be choosen,
NO, because the top index (/) can refer to ANY filename, and if index.html is generally the default name, index.htm can be chosen,
or index.php3, mydog.jpg, or anything you may imagine. (some webmasters are really crazy)
<br>
<br>
@@ -595,8 +554,7 @@ A: <em>Simply use the <tt>--assume dat=application/x-zip</tt> option
A: <em>You may need cookies! Cookies are specific data (for example, your username or password) that are sent to your browser once
you have logged in certain sites so that you only have to log-in once. For example, after having entered your username in a website, you can
view pages and articles, and the next time you will go to this site, you will not have to re-enter your username/password.<br>
To "merge" your personnal cookies to an HTTrack project, just copy the cookies.txt file from your Netscape folder (or the cookies located into the Temporary Internet Files folder for IE)
into your project folder (or even the HTTrack folder)
To supply your own cookies to an HTTrack project, put a Netscape-format <tt>cookies.txt</tt> file in your project folder, or point HTTrack at one with the <tt>--cookies-file</tt> option. You can export your browser's session cookies to that format with a browser extension.
</a><aNAME="Q3b">Q: <strong>I want to update a site, but it's taking too much time! What's happening?</strong><br>
A: <em>First, HTTrack always tries to minimize the download flow by interrogating the server about the
file changes. But, because HTTrack has to rescan all files from the begining to rebuild the local site structure,
file changes. But, because HTTrack has to rescan all files from the beginning to rebuild the local site structure,
it can take some time.
Besides, some servers are not very smart and always consider that they get newer files, forcing HTTrack to reload them,
even if no changes have been made!
@@ -740,7 +698,7 @@ retransferred.</em><br>
</a><aNAME="Q7">Q: <strong>I just want to retrieve all ZIP files or other files in a web
site/in a page. How do I do it?</strong><br>
A: <em>You can use different methods. You can use the 'get files near a link' option if
files are in a foreign domain. You can use, too, a filter adress: adding <tt>+*.zip</tt>
files are in a foreign domain. You can use, too, a filter address: adding <tt>+*.zip</tt>
in the URL list (or in the filter list) will accept all ZIP files, even if these files are
outside the address. <br>
Example : <tt>httrack www.example.com/someaddress.html +*.zip</tt> will allow
@@ -835,8 +793,8 @@ A: <em>Yes. Use user:password@your_proxy_name as your proxy name (example: <tt>s
<br><br><aNAME="QM8">Q: <strong>Can HTTrack generates HP-UX or ISO9660 compatible files?</strong></a><br>
A: <em>Yes. See the build options (-N, or see the WinHTTrack options)</em>
<br><br><aNAME="QM9">Q: <strong>If there any SOCKS support?</strong></a><br>
A: <em>Not yet!</em>
<br><br><aNAME="QM9">Q: <strong>Is there any SOCKS support?</strong></a><br>
A: <em>Yes. HTTrack supports SOCKS5 and HTTP CONNECT proxies: give the proxy with a scheme prefix, e.g. <tt>-P socks5://host:port</tt> or <tt>-P connect://host:port</tt> (prefix <tt>user:pass@</tt> before the host for authenticated proxies).</em>
<br><br><aNAME="QM10">Q: <strong>What's this hts-cache directory? Can I remove it?</strong></a><br>
A: <em>NO if you want to update the site, because this directory is used by HTTrack for this purpose.
<metaname="description"content="HTTrack is an easy-to-use website mirror utility. It allows you to download a World Wide website from the Internet to a local directory,building recursively all structures, getting html, images, and other files from the server to your computer. Links are rebuiltrelatively so that you can freely browse to the local site (works with any browser). You can mirror several sites together so that you can jump from one toanother. You can, also, update an existing mirror site, or resume an interrupted download. The robot is fully configurable, with an integrated help"/>
<metaname="keywords"content="httrack, HTTRACK, HTTrack, winhttrack, WINHTTRACK, WinHTTrack, offline browser, web mirror utility, aspirateur web, surf offline, web capture, www mirror utility, browse offline, local site builder, website mirroring, aspirateur www, internet grabber, capture de site web, internet tool, hors connexion, unix, dos, windows 95, windows 98, solaris, ibm580, AIX 4.0, HTS, HTGet, web aspirator, web aspirateur, libre, GPL, GNU, free software"/>
<metaname="description"content="HTTrack is an easy-to-use website mirror utility. It allows you to download a World Wide website from the Internet to a local directory,building recursively all structures, getting html, images, and other files from the server to your computer. Links are rebuiltrelatively so that you can freely browse to the local site (works with any browser). You can mirror several sites together so that you can jump from one toanother. You can, also, update an existing mirror site, or resume an interrupted download. The robot is fully configurable, with an integrated help"/>
<metaname="keywords"content="httrack, HTTRACK, HTTrack, winhttrack, WINHTTRACK, WinHTTrack, offline browser, web mirror utility, aspirateur web, surf offline, web capture, www mirror utility, browse offline, local site builder, website mirroring, aspirateur www, internet grabber, capture de site web, internet tool, hors connexion, unix, dos, windows 95, windows 98, solaris, ibm580, AIX 4.0, HTS, HTGet, web aspirator, web aspirateur, libre, GPL, GNU, free software"/>
@@ -111,7 +67,7 @@ See also: The <a href="faq.html#VF1">FAQ</a><br>
starts links, the default mode is to mirror these links - i.e. if one of your start page is
www.example.com/test/index.html, all links starting with www.example.com/test/ will be
accepted. But links directly in www.example.com/.. will not be accepted, however, because
they are in a higher strcuture. This prevent HTTrack from mirroring the whole site. (All
they are in a higher structure. This prevent HTTrack from mirroring the whole site. (All
files in structure levels equal or lower than the primary links will be retrieved.)<br>
</i>
<br>
@@ -127,13 +83,15 @@ See also: The <a href="faq.html#VF1">FAQ</a><br>
an authorization filter, like <b><tt>+*.gif</tt></b>. The pattern is a plus (this one: <b><tt>+</tt></b>),
followed by a pattern composed of letters and wildcards (this one: <b><tt>*</tt></b>).
<br><br>
To forbide a family of links, define
To forbid a family of links, define
an authorization filter, like <b><tt>-*.gif</tt></b>. The pattern is a dash (this one: <b><tt>-</tt></b>),
followed by a the same kind of pattern as for the authorization filter.
<br><br>
Example: +*.gif will accept all files finished by .gif<br>
Example: -*.gif will refuse all files finished by .gif<br>
<br>
To see which rule accepted or blocked a given URL, run HTTrack with the <b><tt>--why</tt></b> (<b><tt>-%Y</tt></b>) option, described in <ahref="httrack.man.html#OPTIONS">the manual page</a>.<br>
<br>
<p>
<h4>Scan rules based on size (e.g. accept or refuse files bigger/smaller than a certain size)</h4>
@@ -143,14 +101,14 @@ See also: The <a href="faq.html#VF1">FAQ</a><br>
size to ensure that you won't reach a defined limit.
Example: You may want to accept all files on the domain www.example.com, using '+www.example.com/*',
including gif files inside this domain and outside (eternal images), but not take to large images,
including gif files inside this domain and outside (external images), but not take to large images,
or too small ones (thumbnails)<br>
Excluding gif images smaller than 5KB and images larger than 100KB is therefore a good option;
+www.example.com +*.gif -*.gif*[<5]-*.gif*[>100]
<br>
Important notice: size scan rules are checked <fontcolor=red><b>after</b></font> the link was scheduled for download,
Important notice: size scan rules are checked <b>after</b> the link was scheduled for download,
allowing to abort the connection.
@@ -172,7 +130,7 @@ See also: The <a href="faq.html#VF1">FAQ</a><br>
-mime:image/gif
<br>
Important notice: MIME types scan rules are <fontcolor=red><b>only</b></font> checked against links that were
Important notice: MIME types scan rules are <b>only</b> checked against links that were
scheduled for download, i.e. links <b>already authorized</b> by url scan rules.
Hence, using '+mime:image/gif' will only be a hint to accept images that were already authorized,
if previous MIME scan rules excluded them - such as in '-mime:*/* +mime:text/html +mime:image/gif'
@@ -180,7 +138,7 @@ See also: The <a href="faq.html#VF1">FAQ</a><br>
<h4>1.a. Scan rules based on URL or extension</h4>
@@ -189,7 +147,7 @@ See also: The <a href="faq.html#VF1">FAQ</a><br>
<br>
Filters are analyzed by HTTrack from the first filter to the last one. The complete URL
name is compared to filters defined by the user or added automatically by HTTrack. <br><br>
A scan rule has an higher priority is it is declared later - hierarchy is important: <br>
A scan rule has a higher priority if it is declared later - hierarchy is important: <br>
<br>
<tableBORDER="1"CELLPADDING="2">
@@ -307,7 +265,7 @@ See also: The <a href="faq.html#VF1">FAQ</a><br>
<br>
Filters are analyzed by HTTrack from the first filter to the last one. The sizes
are compared against scan rules defined by the user.<br><br>
A scan rule has an higher priority is it is declared later - hierarchy is important.<br>
A scan rule has a higher priority if it is declared later - hierarchy is important.<br>
Note: scan rules based on size can be mixed with regular URL patterns<br>
@@ -361,7 +319,7 @@ See also: The <a href="faq.html#VF1">FAQ</a><br>
<br>
Filters are analyzed by HTTrack from the first filter to the last one. The complete MIME
type is compared against scan rules defined by the user.<br><br>
A scan rule has an higher priority is it is declared later - hierarchy is important<br>
A scan rule has a higher priority if it is declared later - hierarchy is important<br>
Note: scan rules based on MIME types can <b>NOT</b> be mixed with regular URL patterns or size patterns within the same rule, but you can use both of them in distinct ones<br>
@@ -387,12 +345,12 @@ See also: The <a href="faq.html#VF1">FAQ</a><br>
<tr>
<tdnowrap><tt>-mime:video/*</tt></td>
<td>This will refuse all video links that were already scheduled for download
(i.e. all other 'application/' link download will be aborted)</td>
<metaname="description"content="How to mirror a website with the HTTrack graphical interface: a step-by-step walkthrough and a full option reference for WinHTTrack on Windows, WebHTTrack on Linux and Unix, and HTTrack for Android.">
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.