Compare commits

...

30 Commits

Author SHA1 Message Date
Xavier Roche
6bb609ba27 Merge origin/master into feat/change-report
Conflict in tests/Makefile.am only: master added 88 and 92; the change-report
test keeps 93.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-26 23:28:29 +02:00
Xavier Roche
1d61bf94ca changes: keep the failed-transfer case portable
A connection killed before the status line surfaces differently on macOS, where
it truncates the mirrored file to zero (#748). Assert only what holds on both:
it is never reported gone, and its file survives.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-26 23:23:10 +02:00
Xavier Roche
9484b32ecd An .ndx entry's two URL halves overflow the buffer they share (#744)
Both binput calls pass HTS_URLMAXSIZE, but they write into one
line[HTS_URLMAXSIZE * 2] and binput puts its NUL at s[max]. A first field
that fills its whole bound leaves the second writing line[2048], one past
the array. ASan reports a stack-buffer-overflow on a crafted .ndx. Bound
the second by what the first left.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 23:20:10 +02:00
Xavier Roche
1d647bfecd PT_GetTime's gmtime-failure fallback would print a day of 00 (#742)
* PT_GetTime's failure fallback formats as day 00

An all-zero struct tm has tm_mday == 0, and the ARC filedesc line prints
tm_mday raw, so a gmtime failure would emit "...0100" where the day
belongs. Use the epoch, as PT_SaveCache__Arc_Fun already does for an
unparseable Last-Modified.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* The fallback is reachable, so test it

I claimed no reachable trigger. Wrong: file_timestamp() passes st_mtime
through untouched, and a cache whose mtime is past gmtime's range takes
the fallback. Only the ARC loader is safe, because it overwrites the
timestamp with the filedesc line's own 4-digit-year date.

Test 88 sets an out-of-range mtime on a zip cache and pins the emitted
date; without the fix it reads 19000100000000.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* Register the Windows skip, and make the skip path real

The Windows job pins an exact expected-skip list, so a new conditional
skip fails it; test 88 skips there because NTFS will not hold an mtime
past gmtime's range. Its own skip logic was also dead code: "rc=0 || rc=$?"
never runs the right-hand side, and only set -e was carrying python's 77
through. Capture the status properly and tell a clamp apart from a real
failure.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 23:19:51 +02:00
Xavier Roche
b2dc012263 Route ProxyTrack's remaining raw strcpy through the wrappers (#743)
Fourteen sites, all safe today: literals into sized fields, and two
same-sized array copies. The three writing "" through r->location, a
char *, become a direct terminator, since strcpybuff would have taken its
pointer path and lost the bound.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 22:59:43 +02:00
Xavier Roche
aa1131982b Three hand-written copies of the same clipping contract (#741)
* Fold the three copies of the clip idiom into strclipbuff()

htscache.c and the two readers in proxy/store.c each spelled out
clear-then-strlncatbuff, in binaries that share no code. A helper in
htssafe.h states the contract once, and evaluates its arguments once: the
macro form expanded refvalue and refvalue_size twice, and (refvalue_size)
- 1 would have wrapped to SIZE_MAX had any call site ever passed 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* Pin the capacity-2 boundary

The cases jumped from the degenerate capacity 1 straight to 8, so a
defect confined to small-but-not-degenerate sizes passed: clipping a
two-byte destination to the empty string instead of one character.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 22:58:58 +02:00
Xavier Roche
25ea1d5c4c Merge origin/master into feat/change-report
Conflict in tests/Makefile.am only: master took 89, so the change-report test
moves to 93 (PR #720 holds 92).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-26 22:58:52 +02:00
Xavier Roche
8ebe99be9d changes: prove the fixes, and give the report a cache-off mode
Adds a changes-race self-test (the FTP shape: notifier threads against the
report path), fixtures for a transfer the crawl never completes, for a leftover
file at a name the crawl mirrors fresh, and for a cache-off mirror, plus a pass
with purging on. Registers the web GUI's --changes box with the clearing
companion master's #725 now requires, and documents the degraded mode.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-26 22:57:15 +02:00
Xavier Roche
f72e7ebe96 htsserver: an unauthenticated GET of a directory spins the server forever (#724)
* htsserver: a request naming a directory spins the server forever

fopen() succeeds on a directory on POSIX and every read from it fails with
EISDIR without ever raising EOF, so smallserver()'s "while (!feof(fp))" serving
loop never terminates. GET /server/ needs no session id and no project to reach
it, and the accept loop is single-threaded, so one unauthenticated local request
wedges WebHTTrack for good.

Refuse a directory before fopen() so the 404 branch answers, rather than only
ending the loop: a loop-only fix would serve every directory as an empty 200.
The two loops fed a client-influenced path also stop on ferror(), since a read
that fails for any other reason spins the same way.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* htsserver: whitelist regular files instead of blacklisting directories

fexist() (htsserver.h) is the stat + S_ISREG predicate this file already uses
for the "file-exists:" template op, so gate the serving fopen() on it rather
than on a fresh negated is_directory(). Refusing only directories still handed
FIFOs to the same code path, where fopen() blocks with no writer and kills the
single-threaded accept loop for good.

Test 91 gains the FIFO case and a POST that loads a project whose
hts-cache/winprofile.ini is a directory: that fopen() sits ahead of every guard,
so it covers the ferror() check on the project-load read loop.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* htsserver: fold the two comments above the serve guard into one

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 22:49:59 +02:00
Xavier Roche
a00b95dcb4 changes: fix the review's blocking findings
Lock the report path against the FTP thread the crawl never joins, seal the
accumulator once the report is written, key entries off the project directory
so the report survives --cache=0, stop calling a file gone when the crawl only
failed to re-fetch it, and skip the hook's work entirely when --changes is off.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-26 22:36:08 +02:00
Xavier Roche
3b53bf85f2 A posted project path overflows htsserver's "save settings" error messages (#723)
* htsserver: clip the "save settings" failure messages

The three sprintf(tmp[1024], ...) sites reporting a failed profile save
quote a project path composed from the posted "path" and "projname"
fields, neither of them bounded. A sid-authenticated save with a
1200-byte path smashes the stack buffer and takes the server down.

Route them through a SET_ERRORF() that formats into a fixed buffer and
absorbs the truncation once, so the message clips instead of aborting:
the text comes from the client, and a quieter denial of service is no
better than a loud one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: reach the second message without a 1030-byte mkdir

The tree the case needed came within a few bytes of macOS's PATH_MAX
once /var resolves to /private/var. A symlinked hts-cache gets there
instead: structcheck() lets a non-directory through, and opening
winprofile.ini under it fails with ENOTDIR whatever the uid.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* htssafe: one clipping printf helper for the three failf wrappers

htsblk_failf(), PT_Element_failf() and htsserver's format_error() all had
the same body; slprintfbuff_clip() now owns it. A (void) cast on
slprintfbuff() is no substitute: GCC warns through warn_unused_result.

vslprintfbuff() also empties dest before formatting, so a vsnprintf that
fails outright cannot leave the caller publishing uninitialized stack.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: assert the clipped length, not just the message text

Two of the three sites only checked that the message appeared, which the
old sprintf did equally well; each now compares the rendered message
against the 1023-byte clip. The failing-write branch no longer needs
/dev/full either: the server runs under a file-size limit, so macOS gets
the same coverage.

Renumbered 86 to 89, which no other pending change claims.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: build the oversized profile without a 16k printf width

The macOS leg reached the third message with no error set at all, so the
write the file-size limit was supposed to break had gone through. Build the
profile by doubling a 1024-wide printf, assert its length, and assert the
init file came out short, so a limit that does not bite names itself
instead of surfacing as a missing message.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: build the long paths from the physical temp directory

That is what broke the macOS leg: TMPDIR there resolves through the /var
symlink, so the 1020-byte init file the third case opens was really 1028
bytes to the kernel and the open failed with ENAMETOOLONG. Both cases sit
within a few bytes of macOS's 1024-byte PATH_MAX, so the paths have to be
measured after resolution.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 22:25:52 +02:00
Xavier Roche
fd20086d05 Merge origin/master into feat/change-report
Conflict in tests/Makefile.am only: master took 87 and 90, PR #718 holds 88,
so the change-report test moves to 89.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-26 22:12:02 +02:00
Xavier Roche
ce7dcfa9de Options ticked on by default cannot be un-ticked in the web GUI (#725)
* Options ticked on by default cannot be un-ticked in the web GUI

An unchecked HTML checkbox posts nothing, so htsserver never overwrites the
value it already holds. Every box in the wizard is one-way: once the stored
value is "1", whether htsserver seeded it at startup, a loaded profile set it,
or the user ticked it earlier in the session, un-ticking and submitting leaves
the option on and draws the box ticked again. Only four boxes, all in
option1.html, carried the companion hidden field that guards against this.

Add it to every remaining bare checkbox, and switch cookies and parsejava to
${ztest:...} so a cleared box emits --cookies=0 / --parse-java=0; ${test:...}
renders nothing at all when the value is empty, which is not "off" for an
option the engine turns on by default.

index, urlhack and keep-alive are deliberately left alone: their long options
are declared "single" in htsalias.c and optalias_check drops the =value, so
--index=0 resolves to -I and turns the option back on. That parser bug is
pre-existing and needs its own fix.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: pin last-write-wins and cover every checkbox

The runtime leg posted each name once, so a first-wins body parser would
have passed while the fix did nothing in a real browser: post the
duplicated name in both orders and assert the last value wins. Replace the
four hand-written option cases with a table covering all 28 non-skipped
boxes, asserting the command-line token and the Windows-profile key each
state emits, plus a completeness check so a new box cannot slip through
unexercised.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: rename the loop variable shadowing the scanned page

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 22:08:46 +02:00
Xavier Roche
069573edc3 Test assertions read a padded or truncated reply as a clean security verdict (#728)
* tests: a failed request must not read as a clean security verdict

Under pipefail, "request | grep -q MARKER && fail" skips the fail when the
request itself errors: the leak checks in tests 78 and 85 then pass without
ever having run. Capture the reply first and fail loudly if it never arrived.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* AGENTS.md: record the fail-open assertion shape

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: a reply that proves nothing must not read as a clean verdict

The previous commit converted some of the fail-open assertions and left three.
78's refusal loop still piped into "grep -q ... && fail": grep -q exits on the
first match and SIGPIPEs the producer, so under pipefail a hostile reply that
pads its Location past the 64 KB pipe buffer suppresses the failure exactly as
a dead probe would. 85's fetch() only required a non-empty reply, so a
truncated body or a 302 to the file passed the leak checks marker-free, and no
assertion looked at the status line at all. 78's store probe had no emptiness
guard, so an empty page read as "the store was not written".

Match from here-strings throughout, give fetch() the status each caller
expects, and route 78's store probe through a helper that requires a served
page. 77's X-Injected check had the same shape.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 22:08:10 +02:00
Xavier Roche
a75f437df9 lienrelatif() reads one byte before its stack buffer on an empty path (#729)
The trim that walks back to the last '/' starts at `curr + strlen(curr) - 1`,
which is `curr - 1` when the path is empty. The loop then dereferences it.

An empty path is reachable today: the pre-pass that strips a query does
`strncatbuff(newcurr_fil, curr_fil, a - curr_fil)`, so any `curr_fil` starting
with '?' hands the walk an empty string. `-#test=relative "dir/page.html" "?x"`
under ASan reports the underflow.

The read is one byte and the loop stops immediately either way, so the guard
changes no output: over the 484 ordered pairs of a 22-value path corpus, run
against builds that force the byte before the buffer to 0 and to '/', the
guarded and unguarded results are identical.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 22:06:11 +02:00
Xavier Roche
bcef3a1096 changes: translate the new GUI strings into the remaining 28 locales (#740)
The feature PR added the two LANG_CHANGES* entries to lang.def with English and
Francais only; every other locale fell back to English in the WebHTTrack form.
Each file is written in its own declared charset.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 22:04:50 +02:00
Xavier Roche
783f6ee1f5 AGENTS.md: record what the msg[80] hardening batch taught (#737)
Five PRs across the engine and ProxyTrack turned up the same few traps
more than once, and none of them are obvious from the code.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 21:39:57 +02:00
Xavier Roche
913caf68be Clear the last three compiler warnings (#733)
* Clear the last three compiler warnings

finalurl was sized for one of the two URLs it concatenates. The IIS-bug
example callback overwrites a suffix in place with a same-length
replacement and must not terminate the string, which is memcpy, not
strncpy. The coucal bench's if/else chain has no final else, so result was
only initialised on the paths gcc could not prove exhaustive.

A clean build now reports zero warnings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* Bound the IIS suffix copy by what matched, not by the table

Copying strlen(replacement) leaves the "MUST be the same sizes" comment as
the only thing standing between a future table edit and an overflow. j is
the number of bytes just matched in the destination, so copying j is safe
whatever the table holds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 20:37:52 +02:00
Xavier Roche
59660102d6 A cache field wider than ours aborts the engine instead of clipping (#732)
* A cache field wider than ours aborts the engine instead of clipping

The read-side ZIP_READFIELD_STRING used strlcpybuff, and the whole *_safe_
family aborts on overflow rather than truncating. Since the header line is
bounded only by HTS_URLMAXSIZE and msg is 80 bytes, a cache written by
another build, or a corrupt one, kills the crawl outright. The corrupt-cache
self-test already promises "rejected per-entry, never crash".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* Pin the clip to each field's own capacity

Review found one case exercised only msg[80], so a hardcoded clip length
passed. lastmodified[64] is narrower, and no single constant satisfies
both. The forged replacement was also one byte longer than the line it
overwrote, which only worked because corrupt_patch copies exactly the
pattern length.

Also stop claiming another build's cache can trigger this: the writer
emits each field from the same struct the reader fills, so it takes a
corrupt cache.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 20:37:38 +02:00
Xavier Roche
e96399910b Share the in-progress display struct instead of copying it (#734)
t_StatsBuffer was defined twice, byte for byte, in httrack.h and htsweb.h,
with NStatsBuffer duplicated alongside. httrack and htsserver each keep
their own array, so nothing catches the two drifting apart, and the last
change to the struct had to be applied to both by hand. Both now include
src/htsstats.h; sizes and offsets are unchanged.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 20:19:01 +02:00
Xavier Roche
52d0ab2356 An entry with no usable Last-Modified crashes proxytrack --convert (#731)
* Cached entry with no usable date crashes the ARC writer

PT_SaveCache__Arc_Fun dereferenced convert_time_rfc822() straight into the
record line, so any entry whose Last-Modified is absent or unparseable took
proxytrack --convert down. The sibling caller a thousand lines up already
guards the same call; this one fills in the epoch instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* Assert the archive date, not just the entry

Review found the test blind to the two mutants that matter: a guard firing
unconditionally clobbers every valid date to the epoch, and one that skips
the year and day emits a month and day of 00. Grepping only for the URL saw
neither. Assert the date field exactly, and add a valid-date case so the
untouched path is pinned too.

A bare "Last-Modified: 0" crashes the same way, so it joins the cases.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 20:18:58 +02:00
Xavier Roche
6f4b390aa5 Merge origin/master into feat/change-report
Records the second parent: the previous commit carried master's content but
was committed after a stash cleared MERGE_HEAD, so git never saw it as a merge.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>

# Conflicts:
#	tests/Makefile.am
2026-07-26 19:04:08 +02:00
Xavier Roche
b099e303e5 Merge origin/master into feat/change-report
Both sides appended to tests/Makefile.am's TESTS; kept master's
86_local-proxytrack-cache-longfields.test alongside 88_local-changes.test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
2026-07-26 19:00:24 +02:00
Xavier Roche
6a20dfd998 changes: refresh two stale test comments
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
2026-07-26 18:37:17 +02:00
Xavier Roche
08b32f4878 changes: document the renamed-file case in the format page
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
2026-07-26 18:36:21 +02:00
Xavier Roche
a7ed9c3ca0 changes: fix the review's three findings
The size shortcut compared rendered on-disk sizes even for parsed pages, whose
payload digests describe something else entirely, so it decided the outcome
before the payload comparison could run: a page whose payload never changed but
whose rewritten links moved read as changed. It now only applies when both
digests describe the file on disk.

file_notify() reaches the accumulator from the FTP download thread as well as
the main one, and the lazy allocation, the coucal write and the entries realloc
were all unguarded. Every entry point now takes a mutex, and the HTML hook does
its cache read before taking it, since that read can itself re-enter
file_notify() and move the array.

Two fixtures cover what nothing did: a gzipped direct-to-disk body that changes
at constant length (without the pre-sample before the decoded temp is renamed,
it reads as unchanged), and a page with a fixed payload behind a redirect whose
target is renamed, so its file on disk changes length while its bytes do not.
Both were checked against builds with the respective fix reverted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
2026-07-26 18:35:43 +02:00
Xavier Roche
11c0aa59df changes: drop em dashes from the format page
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
2026-07-26 18:10:20 +02:00
Xavier Roche
f7f428bc93 changes: never leave a stale report behind
A crawl that mirrored nothing created no accumulator, so hts_changes_close_opt
returned without writing and the previous run's report stayed on disk as if it
described this one. Write it whenever --changes is on. The no-data rollback is
the deliberate exception, and is now documented: it restores the previous cache
generation, so leaving the matching report alone is the consistent behaviour.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
2026-07-26 18:08:04 +02:00
Xavier Roche
9061f372e4 Merge origin/master into feat/change-report 2026-07-26 18:03:01 +02:00
Xavier Roche
c38bf69c7d Report what a crawl changed against the previous mirror (--changes)
--update already knows which resources were new, which changed and which the
server called unchanged, and throws it away: the flags reach file_notify() and
go no further than a log line, while deletions exist only as a side effect of
purging. --changes (-%d) keeps all of it and writes hts-changes.json plus a
one-line summary in the log.

"Changed" means the bytes differ, not that the server re-sent the resource.
Comparing the mirrored files directly would not work: HTTrack stamps every
parsed page with the crawl date via the footer, so those bytes differ on every
run. Payloads are compared instead, the previous one coming from the cache for
parsed pages and from the local copy sampled just before it is overwritten for
everything else.

The mirror-relative path, not the URL, is the accumulator's key, so a redirect
and its target that share a save name are one entry; and only the first notify
for a file samples its pre-run state, so a retried transfer is not counted
twice. What counts as already mirrored comes from the previous run's file
index rather than from the file's presence on disk: a partial left by this
crawl's own failed attempt is on disk but was never part of the previous
mirror.

The deleted set is now computed whether or not purging is enabled; unlinking
still happens only under --purge-old.

Closes #714

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
2026-07-26 17:57:07 +02:00
88 changed files with 3196 additions and 197 deletions

View File

@@ -226,8 +226,9 @@ jobs:
# Every gate here exits 77, so an all-skipped suite would report green having
# tested nothing: pin the skips, and floor the passes in case the glob empties.
# footer-overflow skips on Windows (needs a path past MAX_PATH); crange pending #581;
# webdav-mime needs a reapable background listener, which MSYS cannot give it.
expected_skips=" 01_engine-footer-overflow.test 48_local-crange-memresume.test 71_local-crange-repaircache.test 79_local-proxytrack-webdav-mime.test"
# webdav-mime needs a reapable background listener, which MSYS cannot give it;
# badmtime needs a filesystem that stores an mtime past gmtime's range.
expected_skips=" 01_engine-footer-overflow.test 48_local-crange-memresume.test 71_local-crange-repaircache.test 79_local-proxytrack-webdav-mime.test 88_local-proxytrack-badmtime.test"
[ "$pass" -ge 90 ] || { echo "::error::only $pass tests passed ($skip skipped)"; exit 1; }
[ "$skipped" = "$expected_skips" ] || { echo "::error::unexpected skips:$skipped"; exit 1; }
[ "$fail" -eq 0 ] || { echo "::error::failing:$failed"; exit 1; }

View File

@@ -26,6 +26,12 @@ the operational checklist: toolchain, invariants, and how to ship a change.
check`, or `PATH="<bld>/src:$PATH"` for a manual run.
- Give new `.test` scripts `set -e`: the older ones predate the rule, so several
`local-crawl.sh` calls with no `set -e` report PASS on any non-last failure.
- Never assert with `cmd | grep -q MARKER && fail`. Under `pipefail` the
pipeline is non-zero both when `cmd` fails and when `grep -q` matches early
and SIGPIPEs it, so the `&&` never fires and a probe that proved nothing reads
as "marker absent". Capture the reply, assert the status line it must carry
(an empty, truncated or redirected one is marker-free too), then match with a
here-string.
## Hard invariants
- **Generated autotools files are NOT in git.** `configure`, every
@@ -48,6 +54,16 @@ the operational checklist: toolchain, invariants, and how to ship a change.
- Bounds-check every copy. Overflow-safe form: put the untrusted value alone,
`untrusted < limit - controlled` — never `controlled + untrusted < limit`,
which can wrap and pass.
- **Abort or clip is a decision, not a default.** The `*_safe_` helpers
(`strcpybuff`, `strlcpybuff`, `strcatbuff`) **abort** on overflow. Right for
our own data, wrong for anything read back from a cache, a header or the
wire, where it trades a memory smash for a crash on malformed input. Clip
with `dst[0] = '\0'; strlncatbuff(dst, src, size, size - 1)`.
- **A warning class is not the unsafe set.** `-Wformat-truncation` fires only on
a *bounded* `snprintf` whose return is discarded, so an unbounded `sprintf`
into the same buffer never appears on it. Before scoping a hardening pass off
compiler output, grep the unguarded forms yourself (`\bsprintf\s*\(`,
`\bstrcpy\s*\(`, `\bstrcat\s*\(`).
## C conventions
- **Use the `*t` allocator wrappers, never raw libc** (`htssafe.h`):
@@ -84,6 +100,17 @@ Before pushing, and when reviewing others, don't skim for bugs:
layout/ABI, cache/wire format, or a security path? A static or unit check
isn't enough; exercise the wrong behavior at runtime. Claude Code:
`/review-recipe`.
- **Poison a canary, never compare it against zero.** Checking that a
neighbouring field is still `'\0'` cannot see the stray NUL an off-by-one
terminator writes — the exact bug the canary is there for. Fill it with a
non-zero byte, and prove it by killing both the stray-`'X'` and the
stray-NUL mutant. Neither ASan nor `_FORTIFY_SOURCE` sees an overflow that
lands inside the same struct.
- **Overshoot every destination, not one.** A bounds test that oversizes a
single field cannot tell a per-field bound from a one-size-fits-all one, nor
from a fix that bounds that field and leaves its neighbours raw. Exercise
each destination the path touches, spanning at least two capacities, and
check what the code actually emits before writing the expected values.
## Commits
- **Sign-off is mandatory.** Every commit carries a `Signed-off-by` trailer:

263
html/changes.html Normal file
View File

@@ -0,0 +1,263 @@
<html xmlns="http://www.w3.org/1999/xhtml" lang="en">
<head>
<meta http-equiv="Content-Type" content="text/html; charset=iso-8859-1" />
<meta name="description" content="HTTrack is an easy-to-use website mirror utility. It allows you to download a World Wide website from the Internet to a local directory,building recursively all structures, getting html, images, and other files from the server to your computer. Links are rebuiltrelatively so that you can freely browse to the local site (works with any browser). You can mirror several sites together so that you can jump from one toanother. You can, also, update an existing mirror site, or resume an interrupted download. The robot is fully configurable, with an integrated help" />
<meta name="keywords" content="httrack, HTTRACK, HTTrack, winhttrack, WINHTTRACK, WinHTTrack, offline browser, web mirror utility, aspirateur web, surf offline, web capture, www mirror utility, browse offline, local site builder, website mirroring, aspirateur www, internet grabber, capture de site web, internet tool, hors connexion, unix, dos, windows 95, windows 98, solaris, ibm580, AIX 4.0, HTS, HTGet, web aspirator, web aspirateur, libre, GPL, GNU, free software" />
<title>HTTrack Website Copier - Change report format specification</title>
<style type="text/css">
<!--
body {
margin: 0; padding: 0; margin-bottom: 15px; margin-top: 8px;
background: #77b;
}
body, td {
font: 14px "Trebuchet MS", Verdana, Arial, Helvetica, sans-serif;
}
#subTitle {
background: #000; color: #fff; padding: 4px; font-weight: bold;
}
#siteNavigation a, #siteNavigation .current {
font-weight: bold; color: #448;
}
#siteNavigation a:link { text-decoration: none; }
#siteNavigation a:visited { text-decoration: none; }
#siteNavigation .current { background-color: #ccd; }
#siteNavigation a:hover { text-decoration: none; background-color: #fff; color: #000; }
#siteNavigation a:active { text-decoration: none; background-color: #ccc; }
a:link { text-decoration: underline; color: #00f; }
a:visited { text-decoration: underline; color: #000; }
a:hover { text-decoration: underline; color: #c00; }
a:active { text-decoration: underline; }
#pageContent {
clear: both;
border-bottom: 6px solid #000;
padding: 10px; padding-top: 20px;
line-height: 1.65em;
background-image: url(images/bg_rings.gif);
background-repeat: no-repeat;
background-position: top right;
}
#pageContent, #siteNavigation {
background-color: #ccd;
}
.imgLeft { float: left; margin-right: 10px; margin-bottom: 10px; }
.imgRight { float: right; margin-left: 10px; margin-bottom: 10px; }
hr { height: 1px; color: #000; background-color: #000; margin-bottom: 15px; }
h1 { margin: 0; font-weight: bold; font-size: 2em; }
h2 { margin: 0; font-weight: bold; font-size: 1.6em; }
h3 { margin: 0; font-weight: bold; font-size: 1.3em; }
h4 { margin: 0; font-weight: bold; font-size: 1.18em; }
.blak { background-color: #000; }
.hide { display: none; }
.tableWidth { min-width: 400px; }
.tblRegular { border-collapse: collapse; }
.tblRegular td { padding: 6px; background-image: url(fade.gif); border: 2px solid #99c; }
.tblHeaderColor, .tblHeaderColor td { background: #99c; }
.tblNoBorder td { border: 0; }
// -->
</style>
</head>
<table width="76%" border="0" align="center" cellspacing="0" cellpadding="0" class="tableWidth">
<tr>
<td><img src="images/header_title_4.gif" width="400" height="34" alt="HTTrack Website Copier" title="" border="0" id="title" /></td>
</tr>
</table>
<table width="76%" border="0" align="center" cellspacing="0" cellpadding="3" class="tableWidth">
<tr>
<td id="subTitle">Open Source offline browser</td>
</tr>
</table>
<table width="76%" border="0" align="center" cellspacing="0" cellpadding="0" class="tableWidth">
<tr class="blak">
<td>
<table width="100%" border="0" align="center" cellspacing="1" cellpadding="0">
<tr>
<td colspan="6">
<table width="100%" border="0" align="center" cellspacing="0" cellpadding="10">
<tr>
<td id="pageContent">
<!-- ==================== End prologue ==================== -->
<h2 align="center"><em>Change report format specification</em></h2>
<br />
Run with <tt>--changes</tt> (<tt>-%d</tt>), HTTrack writes <tt>hts-changes.json</tt>
in the project directory, next to <tt>hts-log.txt</tt>, describing what the crawl
left new, changed, unchanged and gone compared to the previous mirror. The file is
rewritten from scratch at the end of every run, and the log carries a one-line
summary of the same counts.
<br /><br />
<h3>What "changed" means</h3>
A resource is changed when its bytes differ, not when the server merely re-sent
it. HTTrack compares the payload it just received against the copy the previous
run left behind: for pages it parses, the previous payload comes from the cache
(the file on disk carries the mirror footer and its crawl date, so its bytes
differ on every run); for everything else, the mirrored file is the payload
verbatim and is compared directly.
<br /><br />
Where no digest can be taken on either side, because the cache is disabled or
the previous copy is gone, the report falls back to the transfer signal, and a
server that answers 200 rather than 304 reads as changed. Keeping the cache on
(the default) is what makes the report precise.
<br /><br />
<h3>With the cache off</h3>
<tt>--cache=0</tt> costs the report more than the digest of a parsed page. The
mirror's file index (<tt>hts-cache/new.lst</tt>) is what records which files a
run produced, so without it there is no previous mirror to subtract from: nothing
is reported <tt>gone</tt>, and whether the run is a first crawl cannot be decided
at all, which <tt>first_crawl</tt> states as <tt>null</tt> rather than guess. What
is on disk is still compared byte for byte, so the other three lists stay
meaningful, except for the pages HTTrack parses: those have no cached payload to
compare against and fall back to the transfer signal.
<br /><br />
<h3>Fields</h3>
<ul>
<li><tt>schema</tt>: format version, currently <tt>1</tt>. It is bumped only
on an incompatible change; new fields may appear without one.</li>
<li><tt>generator</tt>: the HTTrack build that wrote the file.</li>
<li><tt>date</tt>: when the report was written, UTC, <tt>YYYY-MM-DDThh:mm:ssZ</tt>.</li>
<li><tt>first_crawl</tt>: true when no index of a previous mirror
(<tt>hts-cache/old.lst</tt>) was found, so there was nothing to compare against and
everything is listed as new. Null when the run kept no index at all and the
question cannot be answered (see above).</li>
<li><tt>partial</tt>: true when the report ran out of memory and lists only
part of the mirror.</li>
<li><tt>purged</tt>: true when <tt>--purge-old</tt> was in effect, so the
files under <tt>gone</tt> were also deleted from disk.</li>
<li><tt>counts</tt>: the size of each of the four lists.</li>
<li><tt>new</tt>, <tt>changed</tt>, <tt>unchanged</tt>, <tt>gone</tt>: the
lists themselves. Every mirrored file appears in exactly one of them.</li>
</ul>
Each entry is an object:
<ul>
<li><tt>url</tt>: the absolute URL the file came from. Empty under
<tt>gone</tt>: deletions are computed from the mirror's file index, which records
paths, not URLs.</li>
<li><tt>file</tt>: the path relative to the mirror root, with forward
slashes. This is the entry's identity: a URL and a redirect that resolve to the
same local file are one entry, not two.</li>
<li><tt>size</tt>: the mirrored file's size in bytes, absent when the file
is not on disk.</li>
<li><tt>previous_size</tt>: under <tt>changed</tt> only, the size of the
copy the previous run left.</li>
</ul>
<br />
<h3>Encoding</h3>
The file is JSON, UTF-8. URLs and local paths reach HTTrack as raw bytes and are
not guaranteed to be valid UTF-8; any byte sequence that is not becomes
U+FFFD (<tt>\ufffd</tt>), so the file always parses. Compare on <tt>file</tt>
rather than on <tt>url</tt> when a mirror is known to carry legacy-charset URLs.
<br /><br />
<h3>Example</h3>
<pre>
{
"schema": 1,
"generator": "HTTrack Website Copier/3.49-14",
"date": "2026-07-26T15:29:03Z",
"first_crawl": false,
"partial": false,
"purged": true,
"counts": { "new": 1, "changed": 1, "unchanged": 1, "gone": 1 },
"new": [
{ "url": "http://example.com/d.html", "file": "example.com/d.html", "size": 280 }
],
"changed": [
{ "url": "http://example.com/a.html", "file": "example.com/a.html", "size": 281, "previous_size": 273 }
],
"unchanged": [
{ "url": "http://example.com/b.html", "file": "example.com/b.html", "size": 277 }
],
"gone": [
{ "url": "", "file": "example.com/c.html" }
]
}
</pre>
<br /><br />
<h3>Notes</h3>
<ul>
<li>A file listed under <tt>gone</tt> is only deleted when <tt>--purge-old</tt> is
on. Left in place it drops out of the mirror's index, so it is reported once and
not again.</li>
<li>A resource whose local file name changed since the previous mirror (a new
MIME type, say) is reported as <tt>new</tt> under its new name; the old name is
reported as <tt>gone</tt> only if the file is still on disk. The two entries are
not paired.</li>
<li>A resource this run tried and failed to transfer also drops out of the
mirror's index, but its previous copy is untouched, so it is reported
<tt>unchanged</tt>. Under <tt>--purge-old</tt> that copy is deleted anyway, and
the report says <tt>gone</tt> to match.</li>
<li>A run that transfers no data at all is rolled back: HTTrack restores the
previous cache generation and leaves the previous report in place, so a lost
connection does not overwrite a good report with an empty one.</li>
<li>Content diffs, and keeping the previous copy of a changed page, are out of
scope: both change what a mirror directory contains.</li>
</ul>
<br /><br />
<!-- ==================== Start epilogue ==================== -->
</td>
</tr>
</table>
</td>
</tr>
</table>
</td>
</tr>
</table>
<table width="76%" border="0" align="center" valign="bottom" cellspacing="0" cellpadding="0">
<tr>
<td id="footer"><small>&copy; 1998-2026 Xavier Roche & other contributors - Web Design: Leto Kauler.</small></td>
</tr>
</table>
</body>
</html>

View File

@@ -482,6 +482,15 @@ add <tt>--warc-cdx</tt> for a sorted CDXJ index, or <tt>--wacz</tt> to bundle th
archive, index and pages into one WACZ for replay tools such as
replayweb.page.</small></p>
<h4>See what a re-crawl changed</h4>
<p><tt>httrack https://example.com/ --update --changes --path mydir</tt><br>
<small>Writes <tt>hts-changes.json</tt> in the project folder, listing every
mirrored file as new, changed, unchanged or gone, plus a one-line summary in the
log. &quot;Changed&quot; means the bytes really differ: a server that answers 200
with the same content it served last time lands in <tt>unchanged</tt>. Deletions
are reported whether or not <tt>--purge-old</tt> is deleting them. The format is
documented in <a href="changes.html">the change report specification</a>.</small></p>
<h4>HTTrack as a fetch tool</h4>
<p><tt>httrack --get https://host/file.bin --path tmp</tt><br>
<small><tt>--get</tt> fetches one file with cache, index, depth, cookies and robots
@@ -490,7 +499,8 @@ all disabled.</small></p>
<p><br>For the complete, always-current option list, see
<a href="httrack.man.html">the manual page</a>. For the filter language, see
<a href="filters.html">filters</a>; for the cache and updates, see
<a href="cache.html">cache</a>.</p>
<a href="cache.html">cache</a>; for the change report, see
<a href="changes.html">changes</a>.</p>
<!-- ==================== Start epilogue ==================== -->
</td>

View File

@@ -129,6 +129,9 @@ The library can be used to write graphical GUIs for httrack, or to run mirrors f
<li><a href="cache.html">Cache format</a></li><br>
HTTrack stores original HTML data and references to downloaded files in a cache, located in the hts-cache directory.
This page describes the HTTrack cache format.
<li><a href="changes.html">Change report format</a></li><br>
With --changes, HTTrack writes hts-changes.json describing what the crawl left new, changed, unchanged and gone
compared to the previous mirror. This page describes that file.
</ul>

View File

@@ -108,16 +108,17 @@ offline browser : copy websites to a local directory</p>
--footer</b> ] [ <b>-%l, --language</b> ] [ <b>-%a,
--accept</b> ] [ <b>-%X, --headers</b> ] [ <b>-C,
--cache[=N]</b> ] [ <b>-k, --store-all-in-cache</b> ] [
<b>-%r, --warc</b> ] [ <b>-%n, --do-not-recatch</b> ] [
<b>-%v, --display</b> ] [ <b>-Q, --do-not-log</b> ] [ <b>-q,
--quiet</b> ] [ <b>-z, --extra-log</b> ] [ <b>-Z,
--debug-log</b> ] [ <b>-v, --verbose</b> ] [ <b>-f,
--file-log</b> ] [ <b>-f2, --single-log</b> ] [ <b>-I,
--index</b> ] [ <b>-%i, --build-top-index</b> ] [ <b>-%I,
--search-index</b> ] [ <b>-pN, --priority[=N]</b> ] [ <b>-S,
--stay-on-same-dir</b> ] [ <b>-D, --can-go-down</b> ] [
<b>-U, --can-go-up</b> ] [ <b>-B, --can-go-up-and-down</b> ]
[ <b>-a, --stay-on-same-address</b> ] [ <b>-d,
<b>-%r, --warc</b> ] [ <b>-%d, --changes</b> ] [ <b>-%n,
--do-not-recatch</b> ] [ <b>-%v, --display</b> ] [ <b>-Q,
--do-not-log</b> ] [ <b>-q, --quiet</b> ] [ <b>-z,
--extra-log</b> ] [ <b>-Z, --debug-log</b> ] [ <b>-v,
--verbose</b> ] [ <b>-f, --file-log</b> ] [ <b>-f2,
--single-log</b> ] [ <b>-I, --index</b> ] [ <b>-%i,
--build-top-index</b> ] [ <b>-%I, --search-index</b> ] [
<b>-pN, --priority[=N]</b> ] [ <b>-S, --stay-on-same-dir</b>
] [ <b>-D, --can-go-down</b> ] [ <b>-U, --can-go-up</b> ] [
<b>-B, --can-go-up-and-down</b> ] [ <b>-a,
--stay-on-same-address</b> ] [ <b>-d,
--stay-on-same-domain</b> ] [ <b>-l, --stay-on-same-tld</b>
] [ <b>-e, --go-everywhere</b> ] [ <b>-%H,
--debug-headers</b> ] [ <b>-%!,
@@ -1143,6 +1144,19 @@ past N bytes, --warc-cdx also writes a sorted CDXJ index,
<td width="4%">
<p>-%d</p></td>
<td width="5%"></td>
<td width="82%">
<p>write hts-changes.json listing what this crawl left new,
changed, unchanged and gone compared to the previous mirror
(--changes)</p> </td></tr>
<tr valign="top" align="left">
<td width="9%"></td>
<td width="4%">
<p>-%n</p></td>
<td width="5%"></td>
<td width="82%">

View File

@@ -98,6 +98,9 @@ ${do:end-if}
<input type="hidden" name="redirect" value="">
<input type="hidden" name="closeme" value="">
<!-- clear if not checked -->
<input type="hidden" name="ftpprox" value="">
${LANG_PROXYTYPE}
<select name="proxytype"
title='${html:LANG_PROXYTYPETIP}' onMouseOver="info('${html:LANG_PROXYTYPETIP}'); return true" onMouseOut="info('&nbsp;'); return true"

View File

@@ -98,6 +98,13 @@ ${do:end-if}
<input type="hidden" name="redirect" value="">
<input type="hidden" name="closeme" value="">
<!-- clear if not checked -->
<input type="hidden" name="errpage" value="">
<input type="hidden" name="external" value="">
<input type="hidden" name="hidepwd" value="">
<input type="hidden" name="hidequery" value="">
<input type="hidden" name="nopurge" value="">
${LANG_I33}
<br>
<select name="build"

View File

@@ -98,6 +98,9 @@ ${do:end-if}
<input type="hidden" name="redirect" value="">
<input type="hidden" name="closeme" value="">
<!-- clear if not checked -->
<input type="hidden" name="windebug" value="">
${LANG_I40c}
<br>

View File

@@ -98,6 +98,11 @@ ${do:end-if}
<input type="hidden" name="redirect" value="">
<input type="hidden" name="closeme" value="">
<!-- clear if not checked -->
<input type="hidden" name="ka" value="">
<input type="hidden" name="remt" value="">
<input type="hidden" name="rems" value="">
<table border="0" width="100%" cellspacing="0">
<tr><td>

View File

@@ -98,6 +98,17 @@ ${do:end-if}
<input type="hidden" name="redirect" value="">
<input type="hidden" name="closeme" value="">
<!-- clear if not checked -->
<input type="hidden" name="cookies" value="">
<input type="hidden" name="parsejava" value="">
<input type="hidden" name="updhack" value="">
<input type="hidden" name="urlhack" value="">
<input type="hidden" name="keepwww" value="">
<input type="hidden" name="keepslashes" value="">
<input type="hidden" name="keepqueryorder" value="">
<input type="hidden" name="toler" value="">
<input type="hidden" name="http10" value="">
<input type="checkbox" name="cookies" ${checked:cookies}
title='${html:LANG_I1b}' onMouseOver="info('${html:LANG_I1b}'); return true" onMouseOut="info('&nbsp;'); return true"
> ${LANG_I58}

View File

@@ -98,6 +98,14 @@ ${do:end-if}
<input type="hidden" name="redirect" value="">
<input type="hidden" name="closeme" value="">
<!-- clear if not checked -->
<input type="hidden" name="warc" value="">
<input type="hidden" name="changes" value="">
<input type="hidden" name="norecatch" value="">
<input type="hidden" name="logf" value="">
<input type="hidden" name="index" value="">
<input type="hidden" name="index2" value="">
<input type="checkbox" name="cache2" ${checked:cache2}
title='${html:LANG_I1e}' onMouseOver="info('${html:LANG_I1e}'); return true" onMouseOut="info('&nbsp;'); return true"
> ${LANG_I61}
@@ -114,6 +122,11 @@ ${LANG_WARCFILE}
>
<br><br>
<input type="checkbox" name="changes" ${checked:changes}
title='${html:LANG_CHANGESTIP}' onMouseOver="info('${html:LANG_CHANGESTIP}'); return true" onMouseOut="info('&nbsp;'); return true"
> ${LANG_CHANGES}
<br><br>
<input type="checkbox" name="norecatch" ${checked:norecatch}
title='${html:LANG_I5b}' onMouseOver="info('${html:LANG_I5b}'); return true" onMouseOut="info('&nbsp;'); return true"
> ${LANG_I34b}

View File

@@ -143,6 +143,7 @@ ${do:copy:StripQuery:stripquery}
${do:copy:StoreAllInCache:cache2}
${do:copy:Warc:warc}
${do:copy:WarcFile:warcfile}
${do:copy:Changes:changes}
${do:copy:LogType:logtype}
${do:copy:UseHTTPProxyForFTP:ftpprox}
${do:copy:ProxyType:proxytype}

View File

@@ -114,7 +114,7 @@ ${do:end-if}
${/* Real commands and ini file generated below */}
<!-- engine commandline -->
<!-- engine commandline; ztest so a cleared default-on option still emits its disabling flag -->
${do:output-mode:html}
<textarea name="command" cols="50" rows="4" style="visibility:hidden">
httrack \
@@ -175,8 +175,8 @@ ${/* -m<n> resets the html limit, so the bare form must precede the -m,<n> one *
\
${unquoted:url2}
\
${test:cookies:--cookies=0:}
${test:parsejava:--parse-java=0:}
${ztest:cookies:--cookies=0:}
${ztest:parsejava:--parse-java=0:}
${test:updhack:--updatehack}
${test:urlhack:--urlhack=0:--urlhack}
${test:keepwww:--keep-www-prefix}
@@ -190,6 +190,7 @@ ${/* -m<n> resets the html limit, so the bare form must precede the -m,<n> one *
${test:cache2:--store-all-in-cache}
${test:warc:--warc}
${test:warcfile:--warc-file "}${arg:warcfile}${test:warcfile:"}
${test:changes:--changes}
${test:norecatch:--do-not-recatch}
${test:logf:--single-log}
${test:logtype:::--extra-log:--debug-log}
@@ -242,6 +243,7 @@ StripQuery=${stripquery}
StoreAllInCache=${ztest:cache2:0:1}
Warc=${ztest:warc:0:1}
WarcFile=${warcfile}
Changes=${ztest:changes:0:1}
LogType=${logtype}
UseHTTPProxyForFTP=${ztest:ftpprox:0:1}
ProxyType=${proxytype}

View File

@@ -1042,3 +1042,7 @@ LANG_WARCFILE
WARC archive name:
LANG_WARCFILETIP
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
LANG_CHANGES
Report what changed since the previous mirror
LANG_CHANGESTIP
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.

View File

@@ -964,3 +964,7 @@ WARC archive name:
Èìå íà WARC àðõèâà:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Íåçàäúëæèòåëíî áàçîâî èìå çà WARC àðõèâà; îñòàâåòå ïðàçíî çà àâòîìàòè÷íî èìåíóâàíå â èçõîäíàòà äèðåêòîðèÿ.
Report what changed since the previous mirror
Îò÷åò çà ïðîìåíèòå ñïðÿìî ïðåäèøíîòî îãëåäàëî
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Çàïèñâàíå è íà hts-changes.json ñúñ ñïèñúê íà íîâèòå, ïðîìåíåíèòå, íåïðîìåíåíèòå è èç÷åçíàëèòå ôàéëîâå ñïðÿìî ïðåäèøíîòî îãëåäàëî.

View File

@@ -964,3 +964,7 @@ WARC archive name:
Nombre del archivo WARC:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Nombre base opcional para el archivo WARC; déjelo en blanco para nombrarlo automáticamente en el directorio de salida.
Report what changed since the previous mirror
Informar de los cambios desde la copia anterior
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Escribir también hts-changes.json con la lista de lo que esta captura deja como nuevo, modificado, sin cambios o desaparecido respecto a la copia anterior.

View File

@@ -964,3 +964,7 @@ WARC archive name:
Název archivu WARC:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Volitelný základní název archivu WARC; ponechte prázdné pro automatické pojmenování ve výstupním adresáøi.
Report what changed since the previous mirror
Hlásit, co se od pøedchozího zrcadlení zmìnilo
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Zapsat také hts-changes.json se seznamem toho, co je oproti pøedchozímu zrcadlení nové, zmìnìné, nezmìnìné nebo chybìjící.

View File

@@ -964,3 +964,7 @@ WARC archive name:
WARC 封存檔名稱:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
WARC 封存檔的選用基本名稱;留空則於輸出目錄中自動命名。
Report what changed since the previous mirror
回報自上次鏡射以來的變更
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
同時寫入 hts-changes.json列出這次擷取相對於上次鏡射的新增、變更、未變更與消失的項目。

View File

@@ -964,3 +964,7 @@ WARC archive name:
WARC 归档名称:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
WARC 归档的可选基本名称;留空则在输出目录中自动命名。
Report what changed since the previous mirror
报告自上次镜像以来的变更
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
同时写入 hts-changes.json列出本次抓取相对于上次镜像的新增、更改、未更改和消失的项目。

View File

@@ -966,3 +966,7 @@ WARC archive name:
Naziv WARC arhive:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Neobavezni osnovni naziv WARC arhive; ostavite prazno za automatsko imenovanje u izlaznom direktoriju.
Report what changed since the previous mirror
Prijavi ¹to se promijenilo od prethodnog zrcaljenja
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Zapi¹i i hts-changes.json s popisom onoga ¹to je u odnosu na prethodno zrcaljenje novo, promijenjeno, nepromijenjeno ili nestalo.

View File

@@ -1012,3 +1012,7 @@ WARC archive name:
Navn på WARC-arkiv:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Valgfrit basisnavn til WARC-arkivet; lad feltet stå tomt for automatisk navngivning i outputmappen.
Report what changed since the previous mirror
Rapportér hvad der er ændret siden den forrige spejling
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Skriv også hts-changes.json med en liste over, hvad denne gennemgang efterlader som nyt, ændret, uændret eller forsvundet i forhold til den forrige spejling.

View File

@@ -964,3 +964,7 @@ WARC archive name:
Name des WARC-Archivs:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Optionaler Basisname für das WARC-Archiv; leer lassen, um es automatisch im Ausgabeverzeichnis zu benennen.
Report what changed since the previous mirror
Melden, was sich seit der vorherigen Spiegelung geändert hat
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Zusätzlich hts-changes.json schreiben, das auflistet, was dieser Durchlauf gegenüber der vorherigen Spiegelung als neu, geändert, unverändert oder verschwunden hinterlässt.

View File

@@ -964,3 +964,7 @@ WARC archive name:
WARC-arhiivi nimi:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
WARC-arhiivi valikuline põhinimi; jäta tühjaks, et see väljundkataloogis automaatselt nimetada.
Report what changed since the previous mirror
Teata, mis on eelmisest peegeldusest saadik muutunud
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Kirjuta ka hts-changes.json, mis loetleb, mis on võrreldes eelmise peegeldusega uus, muutunud, muutumatu või kadunud.

View File

@@ -1012,3 +1012,7 @@ WARC archive name:
WARC archive name:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Report what changed since the previous mirror
Report what changed since the previous mirror
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.

View File

@@ -966,3 +966,7 @@ WARC archive name:
WARC-arkiston nimi:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Valinnainen WARC-arkiston perusnimi; jätä tyhjäksi, jotta se nimetään automaattisesti tulostehakemistoon.
Report what changed since the previous mirror
Raportoi, mikä on muuttunut edellisen peilauksen jälkeen
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Kirjoita myös hts-changes.json, joka luettelee, mikä on edelliseen peilaukseen verrattuna uutta, muuttunutta, muuttumatonta tai kadonnutta.

View File

@@ -1012,3 +1012,7 @@ WARC archive name:
Nom de l'archive WARC :
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Nom de base optionnel pour l'archive WARC ; laissez vide pour le générer automatiquement dans le répertoire de sortie.
Report what changed since the previous mirror
Signaler ce qui a changé depuis le miroir précédent
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Écrire aussi hts-changes.json, qui liste ce que ce crawl laisse nouveau, modifié, inchangé ou disparu par rapport au miroir précédent.

View File

@@ -966,3 +966,7 @@ WARC archive name:
Όνομα αρχείου WARC:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Προαιρετικό βασικό όνομα για το αρχείο WARC. Αφήστε το κενό για αυτόματη ονομασία στον κατάλογο εξόδου.
Report what changed since the previous mirror
Αναφορά των αλλαγών από το προηγούμενο αντίγραφο
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Εγγραφή και του hts-changes.json, που παραθέτει τι είναι νέο, τι άλλαξε, τι έμεινε ίδιο και τι χάθηκε σε σχέση με το προηγούμενο αντίγραφο.

View File

@@ -964,3 +964,7 @@ WARC archive name:
Nome dell'archivio WARC:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Nome di base facoltativo per l'archivio WARC; lascia vuoto per assegnarlo automaticamente nella directory di output.
Report what changed since the previous mirror
Segnala che cosa è cambiato dalla copia precedente
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Scrive anche hts-changes.json, che elenca ciò che questa scansione lascia come nuovo, modificato, invariato o scomparso rispetto alla copia precedente.

View File

@@ -964,3 +964,7 @@ WARC archive name:
WARC アーカイブ名:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
WARC アーカイブの任意のベース名。空欄にすると出力ディレクトリ内で自動的に名前が付けられます。
Report what changed since the previous mirror
前回のミラーからの変更点を報告する
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
hts-changes.json も書き出し、前回のミラーと比べて新規、変更、変更なし、消滅となったものを一覧にします。

View File

@@ -964,3 +964,7 @@ WARC archive name:
¸ÜÕ ÝÐ WARC ÐàåØÒÐâÐ:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
¸×ÑÞàÝÞ ÞáÝÞÒÝÞ ØÜÕ ×Ð WARC ÐàåØÒÐâÐ; ÞáâÐÒÕâÕ ßàÐ×ÝÞ ×Ð ÐÒâÞÜÐâáÚÞ ØÜÕÝãÒÐúÕ ÒÞ Ø×ÛÕ×ÝØÞâ ÔØàÕÚâÞàØãÜ.
Report what changed since the previous mirror
¸×ÒÕáâØ èâÞ Õ ßàÞÜÕÝÕâÞ ÞÔ ßàÕâåÞÔÝÞâÞ ÞÓÛÕÔÐÛÞ
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
·ÐßØèØ Ø hts-changes.json áÞ áߨáÞÚ ÝÐ âÞÐ èâÞ Õ ÝÞÒÞ, ßàÞÜÕÝÕâÞ, ÝÕßàÞÜÕÝÕâÞ ØÛØ ØáçÕ×ÝÐâÞ ÒÞ ÞÔÝÞá ÝÐ ßàÕâåÞÔÝÞâÞ ÞÓÛÕÔÐÛÞ.

View File

@@ -964,3 +964,7 @@ WARC archive name:
WARC archívum neve:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
A WARC archívum opcionális alapneve; hagyja üresen az automatikus elnevezéshez a kimeneti könyvtárban.
Report what changed since the previous mirror
Jelentés arról, mi változott az elõzõ tükrözés óta
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
A hts-changes.json fájl írása is, amely felsorolja, mi új, mi változott, mi maradt változatlan és mi tûnt el az elõzõ tükrözéshez képest.

View File

@@ -964,3 +964,7 @@ WARC archive name:
Naam van WARC-archief:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Optionele basisnaam voor het WARC-archief; laat leeg om het automatisch een naam te geven in de uitvoermap.
Report what changed since the previous mirror
Rapporteren wat er sinds de vorige spiegeling is gewijzigd
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Ook hts-changes.json schrijven met een overzicht van wat deze doorloop nieuw, gewijzigd, ongewijzigd of verdwenen laat ten opzichte van de vorige spiegeling.

View File

@@ -964,3 +964,7 @@ WARC archive name:
Navn på WARC-arkiv:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Valgfritt basisnavn for WARC-arkivet; la feltet stå tomt for automatisk navngivning i utdatamappen.
Report what changed since the previous mirror
Rapporter hva som er endret siden forrige speiling
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Skriv også hts-changes.json som lister opp hva denne gjennomgangen etterlater som nytt, endret, uendret eller forsvunnet sammenlignet med forrige speiling.

View File

@@ -964,3 +964,7 @@ WARC archive name:
Nazwa archiwum WARC:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Opcjonalna nazwa bazowa archiwum WARC; pozostaw puste, aby nazwaæ je automatycznie w katalogu wyj¶ciowym.
Report what changed since the previous mirror
Zg³o¶, co zmieni³o siê od poprzedniej kopii
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Zapisz równie¿ plik hts-changes.json z list± tego, co w porównaniu z poprzedni± kopi± jest nowe, zmienione, niezmienione lub usuniête.

View File

@@ -1012,3 +1012,7 @@ WARC archive name:
Nome do arquivo WARC:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Nome base opcional para o arquivo WARC; deixe em branco para nomeá-lo automaticamente no diretório de saída.
Report what changed since the previous mirror
Relatar o que mudou desde o espelhamento anterior
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Gravar também hts-changes.json listando o que esta captura deixa como novo, alterado, inalterado ou removido em relação ao espelhamento anterior.

View File

@@ -964,3 +964,7 @@ WARC archive name:
Nome do arquivo WARC:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Nome base opcional para o arquivo WARC; deixe em branco para o nomear automaticamente no diretório de saída.
Report what changed since the previous mirror
Comunicar o que mudou desde o espelho anterior
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Escrever também hts-changes.json, que lista o que esta recolha deixa como novo, alterado, inalterado ou desaparecido em relação ao espelho anterior.

View File

@@ -964,3 +964,7 @@ WARC archive name:
Numele arhivei WARC:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Nume de baza optional pentru arhiva WARC; lasati gol pentru a-l denumi automat in directorul de iesire.
Report what changed since the previous mirror
Raporteaza ce s-a schimbat fata de copia anterioara
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Scrie si hts-changes.json, care listeaza ce este nou, modificat, nemodificat sau disparut fata de copia anterioara.

View File

@@ -964,3 +964,7 @@ WARC archive name:
Èìÿ WARC-àðõèâà:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Íåîáÿçàòåëüíîå áàçîâîå èìÿ WARC-àðõèâà; îñòàâüòå ïóñòûì äëÿ àâòîìàòè÷åñêîãî èìåíîâàíèÿ â âûõîäíîì êàòàëîãå.
Report what changed since the previous mirror
Ñîîáùàòü, ÷òî èçìåíèëîñü ñ ïðåäûäóùåãî çåðêàëà
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Òàêæå çàïèñûâàòü hts-changes.json ñî ñïèñêîì òîãî, ÷òî ïî ñðàâíåíèþ ñ ïðåäûäóùèì çåðêàëîì ñòàëî íîâûì, èçìåí¸ííûì, íåèçìåí¸ííûì èëè èñ÷åçëî.

View File

@@ -964,3 +964,7 @@ WARC archive name:
Názov archívu WARC:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Voliteµný základný názov archívu WARC; ponechajte prázdne pre automatické pomenovanie vo výstupnom adresári.
Report what changed since the previous mirror
Oznámi», èo sa od predchádzajúceho zrkadlenia zmenilo
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Zapísa» aj hts-changes.json so zoznamom toho, èo je oproti predchádzajúcemu zrkadleniu nové, zmenené, nezmenené alebo chýbajúce.

View File

@@ -964,3 +964,7 @@ WARC archive name:
Ime arhiva WARC:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Neobvezno osnovno ime arhiva WARC; pustite prazno za samodejno poimenovanje v izhodni mapi.
Report what changed since the previous mirror
Porocaj, kaj se je spremenilo od prejsnjega zrcaljenja
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Zapisi tudi hts-changes.json s seznamom tega, kar je v primerjavi s prejsnjim zrcaljenjem novo, spremenjeno, nespremenjeno ali izginilo.

View File

@@ -964,3 +964,7 @@ WARC archive name:
WARC-arkivets namn:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Valfritt basnamn för WARC-arkivet; lämna tomt för att namnge det automatiskt i utdatakatalogen.
Report what changed since the previous mirror
Rapportera vad som ändrats sedan föregående spegling
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Skriv även hts-changes.json som listar vad den här genomgången lämnar som nytt, ändrat, oförändrat eller försvunnet jämfört med föregående spegling.

View File

@@ -964,3 +964,7 @@ WARC archive name:
WARC arþivi adý:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
WARC arþivi için isteðe baðlý temel ad; çýktý dizininde otomatik adlandýrma için boþ býrakýn.
Report what changed since the previous mirror
Önceki yansýmadan bu yana deðiþenleri bildir
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Ayrýca hts-changes.json yazarak bu taramanýn önceki yansýmaya göre neyi yeni, deðiþmiþ, deðiþmemiþ veya kaybolmuþ býraktýðýný listele.

View File

@@ -964,3 +964,7 @@ WARC archive name:
²ì'ÿ WARC-àðõ³âó:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Íåîáîâ'ÿçêîâà áàçîâà íàçâà WARC-àðõ³âó; çàëèøòå ïîðîæí³ì äëÿ àâòîìàòè÷íîãî íàéìåíóâàííÿ ó âèõ³äíîìó êàòàëîç³.
Report what changed since the previous mirror
Ïîâ³äîìëÿòè, ùî çì³íèëîñÿ ç ïîïåðåäíüîãî äçåðêàëà
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
Òàêîæ çàïèñóâàòè hts-changes.json ç³ ñïèñêîì òîãî, ùî ïîð³âíÿíî ç ïîïåðåäí³ì äçåðêàëîì º íîâèì, çì³íåíèì, íåçì³íåíèì àáî çíèêëèì.

View File

@@ -964,3 +964,7 @@ WARC archive name:
WARC arxivi nomi:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
WARC arxivi uchun ixtiyoriy asosiy nom; chiqish katalogida avtomatik nomlash uchun bo'sh qoldiring.
Report what changed since the previous mirror
Oldingi nusxadan beri nima ozgarganini xabar qilish
Also write hts-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror.
hts-changes.json ham yoziladi: unda ushbu yigish oldingi nusxaga nisbatan nimani yangi, ozgargan, ozgarmagan yoki yoqolgan holda qoldirgani royxati boladi.

View File

@@ -71,7 +71,9 @@ static int mysavename(t_hts_callbackarg * carg, httrackp * opt,
for(j = 0; iisBogus[i][j] == a[j] && iisBogus[i][j] != '\0'; j++) ;
if (iisBogus[i][j] == '\0'
&& (a[j] == '\0' || a[j] == '/' || a[j] == '\\')) {
strncpy(a, iisBogusReplace[i], strlen(iisBogusReplace[i]));
/* j bytes matched, so j fit: copying j cannot overrun whatever the
table holds, and the tail must survive untouched */
memcpy(a, iisBogusReplace[i], (size_t) j);
break;
}
}

View File

@@ -3,7 +3,7 @@
.\"
.\" This file is generated by man/makeman.sh; do not edit by hand.
.\" SPDX-License-Identifier: GPL-3.0-or-later
.TH httrack 1 "23 July 2026" "httrack website copier"
.TH httrack 1 "26 July 2026" "httrack website copier"
.SH NAME
httrack \- offline browser : copy websites to a local directory
.SH SYNOPSIS
@@ -75,6 +75,7 @@ httrack \- offline browser : copy websites to a local directory
[ \fB\-C, \-\-cache[=N]\fR ]
[ \fB\-k, \-\-store\-all\-in\-cache\fR ]
[ \fB\-%r, \-\-warc\fR ]
[ \fB\-%d, \-\-changes\fR ]
[ \fB\-%n, \-\-do\-not\-recatch\fR ]
[ \fB\-%v, \-\-display\fR ]
[ \fB\-Q, \-\-do\-not\-log\fR ]
@@ -279,6 +280,8 @@ create/use a cache for updates and retries (C0 no cache,C1 cache is prioritary,*
store all files in cache (not useful if files on disk) (\-\-store\-all\-in\-cache)
.IP \-%r
write an ISO\-28500 WARC/1.1 archive; \-\-warc\-file NAME sets the output name, \-\-warc\-max\-size N rotates segments past N bytes, \-\-warc\-cdx also writes a sorted CDXJ index, \-\-wacz packages it all as a WACZ file (\-\-warc)
.IP \-%d
write hts\-changes.json listing what this crawl left new, changed, unchanged and gone compared to the previous mirror (\-\-changes)
.IP \-%n
do not re\-download locally erased files (\-\-do\-not\-recatch)
.IP \-%v

View File

@@ -48,7 +48,7 @@ htsserver_LDFLAGS = $(AM_LDFLAGS) $(LDFLAGS_PIE)
lib_LTLIBRARIES = libhttrack.la
htsserver_SOURCES = htsserver.c htsserver.h htsweb.c htsweb.h \
htsserver_SOURCES = htsserver.c htsserver.h htsweb.c htsweb.h htsstats.h \
htscmdline.c htscmdline.h \
htsurlport.c htsurlport.h
proxytrack_SOURCES = proxy/main.c \
@@ -66,7 +66,7 @@ libhttrack_la_SOURCES = htscore.c htsparse.c htsback.c htscache.c \
htscmdline.c htshelp.c htslib.c htsurlport.c htscoremain.c \
htsname.c htsrobots.c htstools.c htswizard.c \
htsalias.c htsthread.c htsindex.c htsbauth.c \
htsmd5.c htscodec.c htswarc.c htsproxy.c htszlib.c htswrap.c htsconcat.c \
htsmd5.c htscodec.c htswarc.c htschanges.c htsproxy.c htszlib.c htswrap.c htsconcat.c \
htsmodules.c htscharset.c punycode.c htsencoding.c htssniff.c \
md5.c \
minizip/ioapi.c minizip/mztools.c minizip/unzip.c minizip/zip.c \
@@ -77,7 +77,7 @@ libhttrack_la_SOURCES = htscore.c htsparse.c htsback.c htscache.c \
htshelp.h htsindex.h htslib.h htsurlport.h htsmd5.h \
htsmodules.h htsname.h htsnet.h htssniff.h \
htsopt.h htsrobots.h htsthread.h \
htstools.h htswizard.h htswrap.h htscodec.h htswarc.h htsproxy.h htszlib.h \
htstools.h htswizard.h htswrap.h htscodec.h htswarc.h htschanges.h htsproxy.h htszlib.h \
htsstrings.h htsarrays.h httrack-library.h \
htscharset.h punycode.h htsencoding.h \
htsentities.h htsentities.sh htsbasiccharsets.sh htscodepages.h \
@@ -87,7 +87,7 @@ libhttrack_la_LIBADD = $(THREADS_LIBS) $(ZLIB_LIBS) $(BROTLI_LIBS) $(ZSTD_LIBS)
libhttrack_la_CFLAGS = $(AM_CFLAGS) -DLIBHTTRACK_EXPORTS -DZLIB_CONST
libhttrack_la_LDFLAGS = $(AM_LDFLAGS) -version-info $(VERSION_INFO)
EXTRA_DIST = httrack.h webhttrack \
EXTRA_DIST = httrack.h htsstats.h webhttrack \
version.rc \
libhttrack.rc \
httrack.rc \

View File

@@ -114,6 +114,8 @@ const char *hts_optalias[][4] = {
"strip [host/pattern=]key1,key2,... from URLs"},
{"cookies-file", "-%K", "param1",
"load extra cookies from a Netscape cookies.txt"},
{"changes", "-%d", "single",
"write hts-changes.json: what this crawl changed vs. the previous mirror"},
{"warc", "-%r", "single", "write an ISO-28500 WARC/1.1 archive of the crawl"},
{"warc-file", "-%rf", "param1", "write a WARC archive to the given base name"},
{"warc-max-size", "-%rs", "param1",

View File

@@ -38,6 +38,7 @@ Please visit our Website: http://www.httrack.com
#include "htsnet.h"
#include "htscore.h"
#include "htswarc.h"
#include "htschanges.h"
#include "htsthread.h"
#include <time.h>
/* END specific definitions */
@@ -728,6 +729,13 @@ int back_finalize(httrackp * opt, cache_back * cache, struct_back * sback,
if ((size = hts_codec_unpack(codec, back[p].tmpfile,
unpacked)) >= 0) {
back[p].r.size = back[p].r.totalsize = size;
if (back[p].r.is_write) {
/* Sample the previous copy now: the rename below replaces
it, and file_notify() only fires once it is gone. */
hts_changes_notify(
opt, back[p].url_adr, back[p].url_fil, back[p].url_sav,
HTS_TRUE, back[p].r.notmodified ? HTS_TRUE : HTS_FALSE);
}
if (!back[p].r.is_write) {
// fichier -> mémoire ; le fichier est écrit plus tard
deleteaddr(&back[p].r);

View File

@@ -142,10 +142,12 @@ struct cache_back_zip_entry {
int compressionMethod;
};
/* A corrupt cache can carry a field wider than ours; clipping it keeps the
entry, where aborting would take the crawl down. */
#define ZIP_READFIELD_STRING(line, value, refline, refvalue, refvalue_size) \
do { \
if (line[0] != '\0' && strfield2(line, refline)) { \
strlcpybuff(refvalue, value, refvalue_size); \
(void) strclipbuff(refvalue, refvalue_size, value); \
line[0] = '\0'; \
} \
} while (0)

View File

@@ -1195,6 +1195,12 @@ int cache_legacy_refused_selftest(httrackp *opt, const char *dir) {
/* --- read-side corruption injection --------------------------------------- */
/* 100 'A's: a placeholder header line long enough to be overwritten by a
forged, over-long X-StatusMessage of the same byte length. */
#define CORRUPT_LONG_ETAG \
"AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA" \
"AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA"
/* canary read back intact after each corruption; victim gets the byte surgery
*/
#define CORRUPT_ADR "corrupt.example.com"
@@ -1219,6 +1225,24 @@ static void corrupt_build(httrackp *opt) {
selftest_close(&cache);
}
/* Like corrupt_build, but the victim carries a 100-char Etag placeholder. */
static void corrupt_build_longetag(httrackp *opt) {
cache_back cache;
memset(corrupt_body_a, 'a', sizeof(corrupt_body_a) - 1);
memset(corrupt_body_b, 'b', sizeof(corrupt_body_b) - 1);
remove(reconcile_st_path(opt, "hts-cache/new.zip"));
remove(reconcile_st_path(opt, "hts-cache/old.zip"));
selftest_open_for_write(&cache, opt);
store_entry(opt, &cache, CORRUPT_ADR, "/canary.html", "canary.html", 200,
"OK", "text/html", "utf-8", "", "", "", "", corrupt_body_a,
strlen(corrupt_body_a));
store_entry(opt, &cache, CORRUPT_ADR, "/victim.html", "victim.html", 200,
"OK", "text/html", "utf-8", "", CORRUPT_LONG_ETAG, "", "",
corrupt_body_b, strlen(corrupt_body_b));
selftest_close(&cache);
}
/* Like corrupt_build, but the victim carries a 20-char Etag whose header line
is later overwritten with a forged oversized X-Size (same byte length). */
static void corrupt_build_etag(httrackp *opt) {
@@ -1403,6 +1427,48 @@ static int corrupt_expect_disk_header(httrackp *opt, LLint wantsize,
return fail;
}
/* An over-long field from a foreign cache must clip, not abort: the entry
still reads, clipped to capacity, and the canary survives. */
static int corrupt_expect_victim_clipped(httrackp *opt, size_t wantmsg,
size_t wantlastmod, const char *what) {
cache_back cache;
htsblk v, c;
char BIGSTK lv[HTS_URLMAXSIZE * 2];
char BIGSTK lc[HTS_URLMAXSIZE * 2];
int fail = 0;
selftest_open_for_read(&cache, opt);
lv[0] = lc[0] = '\0';
v = cache_readex(opt, &cache, CORRUPT_ADR, "/victim.html", "", lv, NULL, 1);
if (v.statuscode != 200) {
fprintf(stderr, "%s: %s: status %d, expected 200\n", selftest_tag, what,
v.statuscode);
fail++;
}
if (wantmsg != (size_t) -1 && strlen(v.msg) != wantmsg) {
fprintf(stderr, "%s: %s: msg len %u, expected %u\n", selftest_tag, what,
(unsigned) strlen(v.msg), (unsigned) wantmsg);
fail++;
}
if (wantlastmod != (size_t) -1 && strlen(v.lastmodified) != wantlastmod) {
fprintf(stderr, "%s: %s: lastmodified len %u, expected %u\n", selftest_tag,
what, (unsigned) strlen(v.lastmodified), (unsigned) wantlastmod);
fail++;
}
c = cache_readex(opt, &cache, CORRUPT_ADR, "/canary.html", "", lc, NULL, 1);
if (c.statuscode != 200) {
fprintf(stderr, "%s: %s: canary tainted (status %d)\n", selftest_tag, what,
c.statuscode);
fail++;
}
if (v.adr != NULL)
freet(v.adr);
if (c.adr != NULL)
freet(c.adr);
selftest_close(&cache);
return fail;
}
/* One zip corruption case: build, patch, then check victim+canary in-session.
*/
static int corrupt_case_zip(httrackp *opt, const char *pat, const char *rep,
@@ -1439,6 +1505,31 @@ int cache_corruption_selftest(httrackp *opt, const char *dir) {
failures += corrupt_expect_victim(opt, "Cache Read Error : Read Data",
"garbled deflate stream");
/* A corrupt cache can hold a field wider than ours. Clipping keeps the
entry; aborting would take the crawl down. Overwrite the placeholder Etag
line in place, same byte length, so the zip offsets stay intact. */
corrupt_build_longetag(opt);
corrupt_patch(opt, "Etag: " CORRUPT_LONG_ETAG, 106,
"X-StatusMessage: " /* 17 + 89 = 106 */
"AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA"
"AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA",
1, 1);
failures +=
corrupt_expect_victim_clipped(opt, sizeof(((htsblk *) 0)->msg) - 1,
(size_t) -1, "over-long X-StatusMessage");
/* lastmodified[64] is narrower than msg[80]: one hardcoded clip length
cannot satisfy both. */
corrupt_build_longetag(opt);
corrupt_patch(opt, "Etag: " CORRUPT_LONG_ETAG, 106,
"Last-Modified: " /* 15 + 91 = 106 */
"AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA"
"AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA",
1, 1);
failures += corrupt_expect_victim_clipped(
opt, (size_t) -1, sizeof(((htsblk *) 0)->lastmodified) - 1,
"over-long Last-Modified");
/* An X-Size above INT_MAX is positive as int64 (slips a bare sign check) but
truncates negative in the (int) cast the malloc uses: a wraparound alloc.
cache_add asserts size fits an int, so such a value only reaches the reader

679
src/htschanges.c Normal file
View File

@@ -0,0 +1,679 @@
/* ------------------------------------------------------------ */
/*
HTTrack Website Copier, Offline Browser for Windows and Unix
Copyright (C) 2026 Xavier Roche and other contributors
SPDX-License-Identifier: GPL-3.0-or-later
This program is free software: you can redistribute it and/or modify
it under the terms of the GNU General Public License as published by
the Free Software Foundation, either version 3 of the License, or
(at your option) any later version.
This program is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
GNU General Public License for more details.
You should have received a copy of the GNU General Public License
along with this program. If not, see <http://www.gnu.org/licenses/>.
Ethical use: we kindly ask that you NOT use this software to harvest email
addresses or to collect any other private information about people. Doing so
would dishonor our work and waste the many hours we have spent on it.
Please visit our Website: http://www.httrack.com
*/
/* ------------------------------------------------------------ */
/* File: htschanges.c subroutines: */
/* --changes: what this crawl changed vs. the previous */
/* mirror (hts-changes.json) */
/* Author: Xavier Roche */
/* ------------------------------------------------------------ */
#define HTS_INTERNAL_BYTECODE
#include "htschanges.h"
#include "htscache.h"
#include "htscharset.h"
#include "htscore.h"
#include "htslib.h"
#include "htsmd5.h"
#include "htssafe.h"
#include "htsthread.h"
#include "htstools.h"
#include "coucal/coucal.h"
#include "md5.h"
#include <stdlib.h>
#include <string.h>
#include <time.h>
#define DIGEST_SIZE 16
typedef struct hts_changes hts_changes;
/* One mirrored file. The mirror-relative path is the key, not the URL: a
redirect and its target can share a save name, and counting that file twice
would put it in two buckets. */
typedef struct {
char *url; /* absolute URL, "" for an engine-generated file */
char *file; /* mirror-relative path, as listed in new.lst */
hts_boolean rewritten; /* the crawl wrote over the local copy */
hts_boolean not_updated; /* transfer signal: the server reported no change */
hts_boolean existed; /* the file was part of the previous mirror */
hts_boolean on_disk; /* a copy was on disk when the crawl first saw it */
hts_boolean listed_prev; /* the previous mirror's index listed this file */
hts_boolean has_prev; /* prev_digest holds the previous payload's digest */
hts_boolean has_new; /* new_digest was taken from the payload, not disk */
unsigned char prev_digest[DIGEST_SIZE];
unsigned char new_digest[DIGEST_SIZE];
LLint prev_size; /* -1 when unknown */
LLint size; /* size after the crawl, -1 when unknown */
hts_change_bucket bucket;
} changes_entry;
struct hts_changes {
coucal index; /* mirror-relative path -> entry slot + 1 */
changes_entry *entries;
size_t count;
size_t capacity;
hts_boolean has_index; /* the run keeps a mirror index (the cache is on) */
hts_boolean old_index; /* the previous mirror's index was read */
hts_boolean overflow; /* an allocation failed; the report is partial */
hts_boolean closed; /* reported: late notifies are dropped, not recorded */
};
/* ------------------------------------------------------------ */
/* JSON */
/* ------------------------------------------------------------ */
/* Length of the UTF-8 sequence led by c, or 0 if c cannot lead one. */
static size_t utf8_lead_length(unsigned char c) {
if (c >= 0xc2 && c <= 0xdf)
return 2;
if (c >= 0xe0 && c <= 0xef)
return 3;
if (c >= 0xf0 && c <= 0xf4)
return 4;
return 0;
}
void hts_changes_json_string(String *out, const char *s) {
const unsigned char *p = (const unsigned char *) s;
StringAddchar(*out, '"');
while (*p != '\0') {
const unsigned char c = *p;
if (c == '"' || c == '\\') {
StringAddchar(*out, '\\');
StringAddchar(*out, (char) c);
p++;
} else if (c < 0x20 || c == 0x7f) {
char esc[8];
snprintf(esc, sizeof(esc), "\\u%04x", (unsigned) c);
StringCat(*out, esc);
p++;
} else if (c < 0x80) {
StringAddchar(*out, (char) c);
p++;
} else {
/* Non-ASCII: emit it verbatim only if it is a well-formed sequence, so
a legacy-charset URL cannot produce unparseable JSON. */
const size_t len = utf8_lead_length(c);
if (len != 0 && strnlen((const char *) p, len) == len &&
hts_isStringUTF8((const char *) p, len)) {
StringMemcat(*out, (const char *) p, len);
p += len;
} else {
StringCat(*out, "\\ufffd");
p++;
}
}
}
StringAddchar(*out, '"');
}
/* ------------------------------------------------------------ */
/* Accumulator */
/* ------------------------------------------------------------ */
static hts_changes *changes_new(void) {
hts_changes *changes = calloct(1, sizeof(*changes));
if (changes == NULL)
return NULL;
changes->index = coucal_new(0);
if (changes->index == NULL) {
freet(changes);
return NULL;
}
return changes;
}
static void changes_free(hts_changes **pchanges) {
hts_changes *changes = *pchanges;
if (changes == NULL)
return;
if (changes->entries != NULL) {
size_t i;
for (i = 0; i < changes->count; i++) {
freet(changes->entries[i].url);
freet(changes->entries[i].file);
}
freet(changes->entries);
}
if (changes->index != NULL)
coucal_delete(&changes->index);
freet(changes);
*pchanges = NULL;
}
/* Slot of `file`, or -1 when it has not been recorded. */
static intptr_t changes_find(const hts_changes *changes, const char *file) {
intptr_t slot = 0;
return coucal_read(changes->index, file, &slot) ? slot - 1 : -1;
}
/* Slot of `file`, appending a fresh entry when it is new; -1 if allocation
failed, which flags the report partial. */
static intptr_t changes_slot(hts_changes *changes, const char *file,
hts_boolean *is_new) {
intptr_t slot = 0;
*is_new = HTS_FALSE;
if (coucal_read(changes->index, file, &slot))
return slot - 1;
if (changes->count == changes->capacity) {
const size_t capacity = changes->capacity != 0 ? changes->capacity * 2 : 64;
changes_entry *const entries =
realloct(changes->entries, capacity * sizeof(*entries));
if (entries == NULL) {
changes->overflow = HTS_TRUE;
return -1;
}
changes->entries = entries;
changes->capacity = capacity;
}
slot = (intptr_t) changes->count;
memset(&changes->entries[slot], 0, sizeof(changes->entries[slot]));
changes->entries[slot].file = strdupt(file);
changes->entries[slot].prev_size = -1;
changes->entries[slot].size = -1;
if (changes->entries[slot].file == NULL) {
changes->overflow = HTS_TRUE;
return -1;
}
changes->count++;
coucal_write(changes->index, file, slot + 1);
*is_new = HTS_TRUE;
return slot;
}
/* MD5 of the file at `path`; not a security boundary. Never call it holding
changes_mutex: hashing a large file would stall every other connection.
HTS_FALSE if it cannot be read. */
static hts_boolean digest_file(const char *path,
unsigned char digest[DIGEST_SIZE]) {
const int endian = 1;
struct MD5Context ctx;
char BIGSTK buffer[32768];
FILE *fp = FOPEN(path, "rb");
size_t nread;
if (fp == NULL)
return HTS_FALSE;
MD5Init(&ctx, *((const char *) &endian));
while ((nread = fread(buffer, 1, sizeof(buffer), fp)) > 0)
MD5Update(&ctx, (const unsigned char *) buffer, (unsigned int) nread);
if (ferror(fp) != 0) {
fclose(fp);
return HTS_FALSE;
}
fclose(fp);
MD5Final(digest, &ctx);
return HTS_TRUE;
}
static void digest_mem(const char *buffer, size_t len,
unsigned char digest[DIGEST_SIZE]) {
domd5mem(buffer, len, (char *) digest, 0);
}
hts_change_bucket hts_changes_classify(hts_boolean rewritten,
hts_boolean existed,
hts_boolean not_updated,
hts_boolean have_digests,
hts_boolean digests_equal) {
if (!rewritten)
return HTS_CHANGE_UNCHANGED; /* the crawl left the copy alone */
if (!existed)
return HTS_CHANGE_NEW;
if (have_digests)
return digests_equal ? HTS_CHANGE_UNCHANGED : HTS_CHANGE_CHANGED;
/* No digest to compare: a server with no validators answers 200 with the
same bytes, so this signal alone over-reports. */
return not_updated ? HTS_CHANGE_UNCHANGED : HTS_CHANGE_CHANGED;
}
/* ------------------------------------------------------------ */
/* Engine hooks */
/* ------------------------------------------------------------ */
/* FTP transfers reach file_notify() from a thread the crawl never joins, so
every entry point below, report and teardown included, takes this lock. */
static htsmutex changes_mutex = HTSMUTEX_INIT;
/* The live accumulator, created on first use; NULL when --changes is off or
the report is already written. Call under the lock. */
static hts_changes *changes_get(httrackp *opt) {
hts_changes *changes = (hts_changes *) opt->changes_state;
if (!opt->changes)
return NULL;
if (changes == NULL) {
changes = changes_new();
opt->changes_state = changes;
}
return changes != NULL && !changes->closed ? changes : NULL;
}
void hts_changes_notify(httrackp *opt, const char *adr, const char *fil,
const char *save, hts_boolean rewritten,
hts_boolean not_updated) {
hts_changes *changes;
char BIGSTK file[HTS_URLMAXSIZE * 2];
char BIGSTK url[HTS_URLMAXSIZE * 4 + 8]; /* holds adr + fil + a scheme */
unsigned char prev_digest[DIGEST_SIZE];
hts_boolean has_prev = HTS_FALSE;
hts_boolean is_new = HTS_FALSE;
intptr_t slot = -1;
LLint prev_size;
if (!opt->changes)
return;
/* Engine-generated scaffolding (the top index) carries no URL and is not a
mirrored resource. */
if (save == NULL || !strnotempty(save) || adr == NULL || fil == NULL ||
(!strnotempty(adr) && !strnotempty(fil)))
return;
hts_savename_listed(StringBuff(opt->path_html), save, file, sizeof(file));
hts_mutexlock(&changes_mutex);
changes = changes_get(opt);
if (changes != NULL) {
slot = changes_slot(changes, file, &is_new);
if (slot >= 0 && !is_new) {
/* Retry, or a second call site for the same file: the pre-run state was
already sampled and the copy on disk may no longer be it. */
changes->entries[slot].rewritten =
changes->entries[slot].rewritten || rewritten;
}
}
hts_mutexrelease(&changes_mutex);
if (slot < 0 || !is_new)
return;
url[0] = '\0';
if (!link_has_authority(adr))
strlcatbuff(url, "http://", sizeof(url));
strlcatbuff(url, adr, sizeof(url));
strlcatbuff(url, fil, sizeof(url));
/* Unlocked: hashing a large file would stall every other connection. Hash
it only when it is about to be overwritten, the last moment it exists. */
prev_size = fsize_utf8(save);
if (prev_size >= 0 && rewritten)
has_prev = digest_file(save, prev_digest);
hts_mutexlock(&changes_mutex);
changes = changes_get(opt);
/* The slot index is stable, entries being append-only; the array is not. */
if (changes != NULL && (size_t) slot < changes->count) {
changes_entry *const entry = &changes->entries[slot];
entry->url = strdupt(url);
if (entry->url == NULL)
changes->overflow = HTS_TRUE;
entry->rewritten = entry->rewritten || rewritten;
entry->not_updated = not_updated;
entry->prev_size = prev_size;
entry->on_disk = prev_size >= 0;
/* hts_changes_html() compares payloads, and its digests win. */
if (has_prev && !entry->has_new) {
entry->has_prev = HTS_TRUE;
memcpy(entry->prev_digest, prev_digest, DIGEST_SIZE);
}
}
hts_mutexrelease(&changes_mutex);
}
void hts_changes_html(httrackp *opt, cache_back *cache, const htsblk *r,
const char *adr, const char *fil, const char *save) {
hts_changes *changes;
char BIGSTK file[HTS_URLMAXSIZE * 2];
char BIGSTK location[HTS_URLMAXSIZE * 2];
unsigned char new_digest[DIGEST_SIZE];
unsigned char prev_digest[DIGEST_SIZE];
hts_boolean has_prev = HTS_FALSE;
intptr_t slot;
htsblk prev;
/* Ahead of the digest and the cache read: off must cost nothing. */
if (!opt->changes)
return;
if (r->adr == NULL || r->size < 0 || save == NULL || !strnotempty(save))
return;
/* On disk this is the payload plus rewritten links and a footer dated by the
crawl, so it differs every run; compare payloads, the previous one being
the body the cache kept. */
digest_mem(r->adr, (size_t) r->size, new_digest);
prev = cache_read_ro(opt, cache, adr, fil, "", location);
if (HTTP_IS_OK(prev.statuscode) && prev.adr != NULL && prev.size >= 0) {
digest_mem(prev.adr, (size_t) prev.size, prev_digest);
has_prev = HTS_TRUE;
}
freet(prev.adr);
/* Only now take the lock and look the entry up: cache_read_ro() can itself
reach file_notify(), which would move the entries array. */
hts_mutexlock(&changes_mutex);
changes = changes_get(opt);
/* file_notify() records the entry; this call only refines it. */
if (changes != NULL) {
hts_savename_listed(StringBuff(opt->path_html), save, file, sizeof(file));
slot = changes_find(changes, file);
if (slot >= 0) {
changes_entry *const entry = &changes->entries[slot];
entry->has_new = HTS_TRUE;
memcpy(entry->new_digest, new_digest, DIGEST_SIZE);
entry->has_prev = has_prev;
if (has_prev)
memcpy(entry->prev_digest, prev_digest, DIGEST_SIZE);
}
}
hts_mutexrelease(&changes_mutex);
}
void hts_changes_dropped(httrackp *opt, const char *file, hts_boolean kept) {
hts_changes *changes;
hts_boolean is_new;
intptr_t slot;
if (!opt->changes || file == NULL || !strnotempty(file))
return;
hts_mutexlock(&changes_mutex);
changes = changes_get(opt);
if (changes != NULL) {
slot = changes_slot(changes, file, &is_new);
if (slot >= 0 && is_new) {
changes->entries[slot].url = strdupt("");
changes->entries[slot].listed_prev = HTS_TRUE;
/* A kept file leaves rewritten clear, which resolves to unchanged. */
if (!kept)
changes->entries[slot].bucket = HTS_CHANGE_GONE;
}
}
hts_mutexrelease(&changes_mutex);
}
void hts_changes_previous(httrackp *opt, const char *file) {
hts_changes *changes;
intptr_t slot;
if (!opt->changes || file == NULL)
return;
hts_mutexlock(&changes_mutex);
changes = changes_get(opt);
if (changes != NULL) {
changes->has_index = HTS_TRUE;
changes->old_index = HTS_TRUE;
slot = changes_find(changes, file);
if (slot >= 0)
changes->entries[slot].listed_prev = HTS_TRUE;
}
hts_mutexrelease(&changes_mutex);
}
void hts_changes_indexed(httrackp *opt) {
hts_changes *changes;
if (!opt->changes)
return;
hts_mutexlock(&changes_mutex);
changes = changes_get(opt);
if (changes != NULL)
changes->has_index = HTS_TRUE;
hts_mutexrelease(&changes_mutex);
}
/* ------------------------------------------------------------ */
/* Report */
/* ------------------------------------------------------------ */
/* Assign each entry its final bucket from the bytes now on disk. */
static void changes_resolve(hts_changes *changes, httrackp *opt) {
char catbuff[CATBUFF_SIZE];
size_t i;
for (i = 0; i < changes->count; i++) {
changes_entry *const entry = &changes->entries[i];
const char *path;
unsigned char digest[DIGEST_SIZE];
hts_boolean have_digests = HTS_FALSE;
hts_boolean digests_equal = HTS_FALSE;
if (entry->bucket == HTS_CHANGE_GONE)
continue;
/* The previous mirror's index is the authority on what was there before:
a partial left by this crawl's own failed attempt is on disk but was
never part of the previous mirror. Without an index, fall back to what
the first notify saw on disk. */
entry->existed = changes->old_index ? entry->listed_prev : entry->on_disk;
path = fconcat(catbuff, sizeof(catbuff), StringBuff(opt->path_html),
entry->file);
entry->size = fsize_utf8(path);
if (entry->rewritten && entry->existed) {
/* The size shortcut only holds when both digests describe the file on
disk; has_new means they are payload digests, and a parsed page's
rendered size moves with its links and footer. */
if (!entry->has_new && entry->size >= 0 && entry->prev_size >= 0 &&
entry->size != entry->prev_size) {
have_digests = HTS_TRUE; /* different lengths: no need to hash */
digests_equal = HTS_FALSE;
} else if (entry->has_prev &&
(entry->has_new
? (memcpy(digest, entry->new_digest, DIGEST_SIZE), 1)
: digest_file(path, digest))) {
have_digests = HTS_TRUE;
digests_equal = memcmp(digest, entry->prev_digest, DIGEST_SIZE) == 0
? HTS_TRUE
: HTS_FALSE;
}
}
entry->bucket =
hts_changes_classify(entry->rewritten, entry->existed,
entry->not_updated, have_digests, digests_equal);
}
}
static const char *const bucket_names[HTS_CHANGE_BUCKETS] = {
"new", "changed", "unchanged", "gone"};
/* Serialize the report. Call under changes_mutex. */
static void changes_serialize(httrackp *opt, String *out) {
hts_changes *const changes = (hts_changes *) opt->changes_state;
size_t counts[HTS_CHANGE_BUCKETS];
char date[32];
char scratch[64];
int bucket;
size_t i;
StringClear(*out);
if (changes == NULL)
return;
changes_resolve(changes, opt);
memset(counts, 0, sizeof(counts));
for (i = 0; i < changes->count; i++)
counts[changes->entries[i].bucket]++;
hts_now_iso8601(date);
StringCat(*out, "{\n \"schema\": ");
snprintf(scratch, sizeof(scratch), "%d", HTS_CHANGES_SCHEMA);
StringCat(*out, scratch);
StringCat(*out, ",\n \"generator\": ");
hts_changes_json_string(out, "HTTrack Website Copier/" HTTRACK_VERSION);
StringCat(*out, ",\n \"date\": ");
hts_changes_json_string(out, date);
StringCat(*out, ",\n \"first_crawl\": ");
StringCat(*out, !changes->has_index
? "null"
: (changes->old_index ? "false" : "true"));
StringCat(*out, ",\n \"partial\": ");
StringCat(*out, changes->overflow ? "true" : "false");
StringCat(*out, ",\n \"purged\": ");
StringCat(*out, opt->delete_old ? "true" : "false");
StringCat(*out, ",\n \"counts\": {");
for (bucket = 0; bucket < HTS_CHANGE_BUCKETS; bucket++) {
StringCat(*out, bucket != 0 ? ", " : " ");
hts_changes_json_string(out, bucket_names[bucket]);
snprintf(scratch, sizeof(scratch), ": %d", (int) counts[bucket]);
StringCat(*out, scratch);
}
StringCat(*out, " }");
for (bucket = 0; bucket < HTS_CHANGE_BUCKETS; bucket++) {
hts_boolean first = HTS_TRUE;
StringCat(*out, ",\n ");
hts_changes_json_string(out, bucket_names[bucket]);
StringCat(*out, ": [");
for (i = 0; i < changes->count; i++) {
const changes_entry *const entry = &changes->entries[i];
if (entry->bucket != bucket)
continue;
StringCat(*out, first ? "\n { \"url\": " : ",\n { \"url\": ");
first = HTS_FALSE;
hts_changes_json_string(out, entry->url != NULL ? entry->url : "");
StringCat(*out, ", \"file\": ");
hts_changes_json_string(out, entry->file);
if (entry->size >= 0) {
snprintf(scratch, sizeof(scratch), ", \"size\": " LLintP,
(LLint) entry->size);
StringCat(*out, scratch);
}
if (bucket == HTS_CHANGE_CHANGED && entry->prev_size >= 0) {
snprintf(scratch, sizeof(scratch), ", \"previous_size\": " LLintP,
(LLint) entry->prev_size);
StringCat(*out, scratch);
}
StringCat(*out, " }");
}
StringCat(*out, first ? "]" : "\n ]");
}
StringCat(*out, "\n}\n");
}
void hts_changes_report(httrackp *opt, String *out) {
hts_mutexlock(&changes_mutex);
changes_serialize(opt, out);
hts_mutexrelease(&changes_mutex);
}
void hts_changes_close_opt(httrackp *opt) {
char catbuff[CATBUFF_SIZE];
const char *path;
String report = STRING_EMPTY;
size_t counts[HTS_CHANGE_BUCKETS];
size_t total;
hts_boolean has_index, old_index;
FILE *fp;
hts_mutexlock(&changes_mutex);
{
/* changes_get(), not the raw field: a crawl that mirrored nothing still
owes the user a report, and a stale one from the previous run must not
survive on disk as if it described this one. */
hts_changes *const changes = changes_get(opt);
size_t i;
if (changes == NULL) {
hts_mutexrelease(&changes_mutex);
return;
}
changes_serialize(opt, &report);
memset(counts, 0, sizeof(counts));
for (i = 0; i < changes->count; i++)
counts[changes->entries[i].bucket]++;
total = changes->count;
has_index = changes->has_index;
old_index = changes->old_index;
/* Sticky: whatever the crawl's stragglers do next is not in this report. */
changes->closed = HTS_TRUE;
}
hts_mutexrelease(&changes_mutex);
path = fconcat(catbuff, sizeof(catbuff), StringBuff(opt->path_log),
HTS_CHANGES_FILE);
fp = FOPEN(path, "wb");
if (fp != NULL) {
const size_t len = StringLength(report);
if (len != 0 && fwrite(StringBuff(report), 1, len, fp) != len)
hts_log_print(opt, LOG_ERROR | LOG_ERRNO,
"Unable to write the change report %s", path);
fclose(fp);
} else {
hts_log_print(opt, LOG_ERROR | LOG_ERRNO,
"Unable to create the change report %s", path);
}
if (!has_index) {
hts_log_print(opt, LOG_NOTICE,
"Change report: no mirror index (the cache is off), %d files "
"mirrored, deletions not detected (%s)",
(int) total, HTS_CHANGES_FILE);
} else if (!old_index) {
hts_log_print(opt, LOG_NOTICE,
"Change report: first crawl, %d files mirrored, nothing to "
"compare against (%s)",
(int) total, HTS_CHANGES_FILE);
} else {
hts_log_print(opt, LOG_NOTICE,
"Change report: %d new, %d changed, %d unchanged, %d gone "
"(%s)",
(int) counts[HTS_CHANGE_NEW],
(int) counts[HTS_CHANGE_CHANGED],
(int) counts[HTS_CHANGE_UNCHANGED],
(int) counts[HTS_CHANGE_GONE], HTS_CHANGES_FILE);
}
StringFree(report);
}
void hts_changes_free_opt(httrackp *opt) {
hts_changes *changes;
hts_mutexlock(&changes_mutex);
changes = (hts_changes *) opt->changes_state;
changes_free(&changes);
opt->changes_state = NULL;
hts_mutexrelease(&changes_mutex);
}

125
src/htschanges.h Normal file
View File

@@ -0,0 +1,125 @@
/* ------------------------------------------------------------ */
/*
HTTrack Website Copier, Offline Browser for Windows and Unix
Copyright (C) 2026 Xavier Roche and other contributors
SPDX-License-Identifier: GPL-3.0-or-later
This program is free software: you can redistribute it and/or modify
it under the terms of the GNU General Public License as published by
the Free Software Foundation, either version 3 of the License, or
(at your option) any later version.
This program is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
GNU General Public License for more details.
You should have received a copy of the GNU General Public License
along with this program. If not, see <http://www.gnu.org/licenses/>.
Ethical use: we kindly ask that you NOT use this software to harvest email
addresses or to collect any other private information about people. Doing so
would dishonor our work and waste the many hours we have spent on it.
Please visit our Website: http://www.httrack.com
*/
/* ------------------------------------------------------------ */
/* HTTrack change report (--changes). Internal, not installed.
Accumulates what the crawl did to each mirrored file, compares the bytes
against the copy the previous run left behind, and writes
hts-changes.json next to the log. */
/* ------------------------------------------------------------ */
#ifndef HTS_CHANGES_DEFH
#define HTS_CHANGES_DEFH
#include "htsopt.h"
#include "htsstrings.h"
#ifndef HTS_DEF_FWSTRUCT_cache_back
#define HTS_DEF_FWSTRUCT_cache_back
typedef struct cache_back cache_back;
#endif
#ifdef __cplusplus
extern "C" {
#endif
/* Report file name, written under the project's log directory. */
#define HTS_CHANGES_FILE "hts-changes.json"
/* Schema version carried by the report; bump on an incompatible change. */
#define HTS_CHANGES_SCHEMA 1
/* Which side of the comparison a mirrored file ended up on. */
typedef enum {
HTS_CHANGE_NEW = 0, /* no local copy before this run */
HTS_CHANGE_CHANGED, /* rewritten, and the bytes differ */
HTS_CHANGE_UNCHANGED, /* the bytes are those of the previous mirror */
HTS_CHANGE_GONE, /* in the previous mirror, absent from this one */
HTS_CHANGE_BUCKETS
} hts_change_bucket;
/* Record what this crawl is doing to the local file `save` (absolute path;
adr/fil form its URL, either may be empty for an engine-generated file).
`rewritten` means the copy on disk is being written over, `not_updated` is
the 200-versus-304 signal, used only when no digest can be taken. Only the
first call for a file samples the previous copy, so a retried or
twice-notified resource is counted once. No-op unless --changes is on. */
void hts_changes_notify(httrackp *opt, const char *adr, const char *fil,
const char *save, hts_boolean rewritten,
hts_boolean not_updated);
/* Refine the entry for a parsed HTML file: its local copy carries rewritten
links and a footer dated by the crawl, so it differs every run. Compares the
payload `r` holds against the body the previous run left in the cache. */
void hts_changes_html(httrackp *opt, cache_back *cache, const htsblk *r,
const char *adr, const char *fil, const char *save);
/* Record `file` (mirror-relative, as listed in new.lst) as listed by the
previous mirror's index and absent from this run's. `kept` means its local
copy survives the run because the crawl tried and failed to replace it: a
failed transfer is not a deletion, so it is reported unchanged, not gone. */
void hts_changes_dropped(httrackp *opt, const char *file, hts_boolean kept);
/* Record that the previous mirror's index listed `file`. That index, not the
file's presence on disk, decides what counts as already mirrored. */
void hts_changes_previous(httrackp *opt, const char *file);
/* Record that this run keeps a mirror index. Without one (the cache is off)
deletions and first_crawl are undecidable, and the report says so. */
void hts_changes_indexed(httrackp *opt);
/* Resolve every entry against the bytes now on disk and serialize the report
into `out` (replaced). Exposed for the self-tests. */
void hts_changes_report(httrackp *opt, String *out);
/* Write the report and log a one-line summary. Idempotent, and seals the
accumulator: a late notify is dropped rather than starting a second one. */
void hts_changes_close_opt(httrackp *opt);
/* Drop the accumulator, for a run that never reached its end and to start the
next one clean. Null-safe and idempotent. */
void hts_changes_free_opt(httrackp *opt);
/* Bucket for one entry, from what was observed. No filesystem access:
`have_digests` says both sides could be hashed, `digests_equal` compares
them, and `not_updated` is the fallback signal when they could not. */
hts_change_bucket hts_changes_classify(hts_boolean rewritten,
hts_boolean existed,
hts_boolean not_updated,
hts_boolean have_digests,
hts_boolean digests_equal);
/* Append `s` to `out` as a quoted JSON string. Byte sequences that are not
valid UTF-8 become U+FFFD, so a mirror carrying legacy-charset URLs still
produces parseable JSON. */
void hts_changes_json_string(String *out, const char *s);
#ifdef __cplusplus
}
#endif
#endif

View File

@@ -40,6 +40,7 @@ Please visit our Website: http://www.httrack.com
/* File defs */
#include "htscore.h"
#include "htswarc.h"
#include "htschanges.h"
/* specific definitions */
#include "htsbase.h"
@@ -488,6 +489,7 @@ void hts_finish_html_file(httrackp *opt, cache_back *cache, htsblk *r,
const char *adr, const char *fil, const char *save) {
{
file_notify(opt, adr, fil, save, 1, 1, r->notmodified);
hts_changes_html(opt, cache, r, adr, fil, save);
*fp = filecreate(&opt->state.strc, save);
if (*fp) {
if (ht_len > 0 && fwrite(ht_buff, 1, ht_len, *fp) != ht_len) {
@@ -691,6 +693,9 @@ int httpmirror(char *url1, httrackp * opt) {
// hash table
opt->hash = &hash;
// a change report left by a previous crawl on this opt is not this one's
hts_changes_free_opt(opt);
// initialize link heap
hts_record_init(opt);
@@ -2089,9 +2094,12 @@ int httpmirror(char *url1, httrackp * opt) {
if (cache.lst) {
fclose(cache.lst);
cache.lst = opt->state.strc.lst = NULL;
if (opt->delete_old) {
/* old.lst minus new.lst is the set of files the previous mirror had and
this one does not. --changes reports it; only --purge-old acts on it. */
if (opt->delete_old || opt->changes) {
FILE *old_lst, *new_lst;
hts_changes_indexed(opt);
//
opt->state._hts_in_html_parsing = 3;
//
@@ -2117,21 +2125,37 @@ int httpmirror(char *url1, httrackp * opt) {
int purge = 0;
while(!feof(old_lst)) {
linput(old_lst, line, 1000);
if (!strstr(adr, line)) { // not found in the new list?
char BIGSTK file[HTS_URLMAXSIZE * 2];
char BIGSTK file[HTS_URLMAXSIZE * 2];
strcpybuff(file, StringBuff(opt->path_html));
strcatbuff(file, line + 1);
file[strlen(file) - 1] = '\0';
if (fexist_utf8(file)) { // toujours sur disque: virer
hts_log_print(opt, LOG_INFO, "Purging %s", file);
UNLINK(file);
purge = 1;
linput(old_lst, line, 1000);
if (!strnotempty(line))
continue;
strcpybuff(file, StringBuff(opt->path_html));
strcatbuff(file, line + 1);
file[strlen(file) - 1] = '\0';
hts_changes_previous(opt, file + StringLength(opt->path_html));
if (!strstr(adr, line)) { // not found in the new list?
if (fexist_utf8(file)) { // still on disk
/* A link this crawl did try but never wrote (a transfer
killed mid-flight) also drops out of new.lst. Unless it
is about to be purged, its previous copy stands and the
file is not gone. */
const hts_boolean kept =
!opt->delete_old &&
hash_read(opt->hash, file, NULL,
HASH_STRUCT_FILENAME) >= 0;
hts_changes_dropped(
opt, file + StringLength(opt->path_html), kept);
if (opt->delete_old) {
hts_log_print(opt, LOG_INFO, "Purging %s", file);
UNLINK(file);
purge = 1;
}
}
}
}
{
if (opt->delete_old) { // emptied directories go with the files
fseek(old_lst, 0, SEEK_SET);
while(!feof(old_lst)) {
linput(old_lst, line, 1000);
@@ -2166,7 +2190,7 @@ int httpmirror(char *url1, httrackp * opt) {
}
}
//
if (!purge) {
if (opt->delete_old && !purge) {
hts_log_print(opt, LOG_INFO, "No files purged");
}
}
@@ -2256,6 +2280,7 @@ int httpmirror(char *url1, httrackp * opt) {
// ending
usercommand(opt, 0, NULL, NULL, NULL, NULL);
warc_close_opt(opt);
hts_changes_close_opt(opt);
// désallocation mémoire & buffers
XH_uninit;
@@ -2881,6 +2906,17 @@ int filecreateempty(filenote_strc * strc, const char *filename) {
return 0;
}
void hts_savename_listed(const char *root, const char *s, char *dest,
size_t destsize) {
char catbuff[CATBUFF_SIZE];
strlcpybuff(dest, fslash(catbuff, sizeof(catbuff), s), destsize);
if (strnotempty(root) && strncmp(fslash(catbuff, sizeof(catbuff), root), dest,
strlen(root)) == 0) {
strlcpybuff(dest, s + strlen(root), destsize);
}
}
// noter fichier
int filenote(filenote_strc * strc, const char *s, filecreate_params * params) {
// gestion du fichier liste liste
@@ -2890,15 +2926,8 @@ int filenote(filenote_strc * strc, const char *s, filecreate_params * params) {
return 0;
} else if (strc->lst) {
char BIGSTK savelst[HTS_URLMAXSIZE * 2];
char catbuff[CATBUFF_SIZE];
strcpybuff(savelst, fslash(catbuff, sizeof(catbuff), s));
// couper chemin?
if (strnotempty(strc->path)) {
if (strncmp(fslash(catbuff, sizeof(catbuff), strc->path), savelst, strlen(strc->path)) == 0) { // couper
strcpybuff(savelst, s + strlen(strc->path));
}
}
hts_savename_listed(strc->path, s, savelst, sizeof(savelst));
fprintf(strc->lst, "[%s]" LF, savelst);
fflush(strc->lst);
}
@@ -2908,6 +2937,9 @@ int filenote(filenote_strc * strc, const char *s, filecreate_params * params) {
/* Note: utf-8 */
void file_notify(httrackp * opt, const char *adr, const char *fil,
const char *save, int create, int modify, int not_updated) {
hts_changes_notify(opt, adr, fil, save,
(create || modify) ? HTS_TRUE : HTS_FALSE,
not_updated ? HTS_TRUE : HTS_FALSE);
RUN_CALLBACK6(opt, filesave2, adr, fil, save, create, modify, not_updated);
}
@@ -3642,6 +3674,7 @@ HTSEXT_API int copy_htsopt(const httrackp * from, httrackp * to) {
to->warc_max_size = from->warc_max_size;
to->warc_cdx = from->warc_cdx;
to->warc_wacz = from->warc_wacz;
to->changes = from->changes;
if (from->pause_max_ms > 0) {
to->pause_min_ms = from->pause_min_ms;

View File

@@ -354,6 +354,12 @@ int filecreateempty(filenote_strc * strct, const char *filename);
int filenote(filenote_strc * strct, const char *s, filecreate_params * params);
/* Copy into dest (destsize bytes) the form under which the local path s is
listed in new.lst: forward slashes, with the mirror `root` stripped when s
sits under it. Also the change report's key, so the two must not drift. */
void hts_savename_listed(const char *root, const char *s, char *dest,
size_t destsize);
void file_notify(httrackp * opt, const char *adr, const char *fil,
const char *save, int create, int modify, int wasupdated);

View File

@@ -41,6 +41,7 @@ Please visit our Website: http://www.httrack.com
#include "htsdefines.h"
#include "htsalias.h"
#include "htswarc.h"
#include "htschanges.h"
#include "htsbauth.h"
#include "htswrap.h"
#include "htsmodules.h"
@@ -1747,6 +1748,13 @@ static int hts_main_internal(int argc, char **argv, httrackp * opt) {
StringCopy(opt->cookies_file, argv[na]);
}
break;
case 'd': // --changes: report what this crawl changed
opt->changes = HTS_TRUE;
if (*(com + 1) == '0') {
opt->changes = HTS_FALSE;
com++;
}
break;
case 'r': // warc / warc-file: write an ISO-28500 WARC archive
if (*(com + 1) == 'f') { // --warc-file NAME: explicit basename
com++;

View File

@@ -433,7 +433,8 @@ void help_catchurl(const char *dest_path) {
}
// former URL!
{
char BIGSTK finalurl[HTS_URLMAXSIZE * 2];
/* url and dest are each HTS_URLMAXSIZE*2, plus the POSTTOK marker */
char BIGSTK finalurl[HTS_URLMAXSIZE * 4 + 32];
inplace_escape_check_url(dest, sizeof(dest));
snprintf(finalurl, sizeof(finalurl), "%s" POSTTOK "file:%s", url, dest);
@@ -607,6 +608,8 @@ void help(const char *app, int more) {
"output name, --warc-max-size N rotates segments past N bytes, "
"--warc-cdx also writes a sorted CDXJ index, --wacz packages it all "
"as a WACZ file");
infomsg(" %d write hts-changes.json listing what this crawl left new, "
"changed, unchanged and gone compared to the previous mirror");
infomsg(" %n do not re-download locally erased files");
infomsg
(" %v display on screen filenames downloaded (in realtime) - * %v1 short version - %v2 full animation");

View File

@@ -37,6 +37,7 @@ Please visit our Website: http://www.httrack.com
#include "htscore.h"
#include "htswarc.h"
#include "htschanges.h"
/* specific definitions */
#include "htsbase.h"
@@ -2537,15 +2538,6 @@ void fil_simplifie(char *f) {
}
}
void htsblk_failf(htsblk *r, const char *fmt, ...) {
va_list args;
va_start(args, fmt);
// deliberate clip: the reason is quoted from a remote reply
(void) vslprintfbuff(r->msg, sizeof(r->msg), fmt, args);
va_end(args);
}
// fermer liaison fichier ou socket
void deletehttp(htsblk * r) {
#if HTS_DEBUG_CLOSESOCK
@@ -2707,6 +2699,24 @@ void time_gmt_rfc822(char *s) {
time_rfc822(s, A);
}
void hts_now_iso8601(char out[32]) {
time_t t = time(NULL);
struct tm tmv;
#if defined(_WIN32)
struct tm *g = gmtime(&t);
if (g != NULL)
tmv = *g;
else
memset(&tmv, 0, sizeof(tmv));
#else
if (gmtime_r(&t, &tmv) == NULL)
memset(&tmv, 0, sizeof(tmv));
#endif
strftime(out, 32, "%Y-%m-%dT%H:%M:%SZ", &tmv);
}
// heure actuelle, format rfc (taille buffer 256o)
void time_local_rfc822(char *s) {
time_t tt;
@@ -6020,6 +6030,8 @@ HTSEXT_API httrackp *hts_create_opt(void) {
StringCopy(opt->cookies_file, "");
StringCopy(opt->warc_file, "");
opt->warc_max_size = 0; /* no rotation unless --warc-max-size sets it */
opt->changes = HTS_FALSE;
opt->changes_state = NULL;
StringCopy(opt->why_url, "");
opt->pause_min_ms = 0;
opt->pause_max_ms = 0;
@@ -6173,6 +6185,8 @@ HTSEXT_API void hts_free_opt(httrackp * opt) {
StringFree(opt->why_url);
StringFree(opt->warc_file);
hts_changes_free_opt(opt);
StringFree(opt->path_html);
StringFree(opt->path_html_utf8);
StringFree(opt->path_log);

View File

@@ -201,7 +201,8 @@ T_SOC newhttp_addr(httrackp *opt, const char *iadr, htsblk *retour, int port,
int waitconnect, int addr_index, int *addr_count);
/* Clips the formatted failure reason into r->msg, which also round-trips
through the cache as X-StatusMessage. Leaves r->statuscode to the caller. */
void htsblk_failf(htsblk *r, const char *fmt, ...) HTS_PRINTF_FUN(2, 3);
#define htsblk_failf(R, ...) \
slprintfbuff_clip((R)->msg, sizeof((R)->msg), __VA_ARGS__)
HTS_INLINE void deletehttp(htsblk * r);
HTS_INLINE int deleteaddr(htsblk * r);
@@ -257,6 +258,10 @@ void sec2str(char *s, TStamp t);
void time_gmt_rfc822(char *s);
void time_local_rfc822(char *s);
/* Current UTC time as "YYYY-MM-DDThh:mm:ssZ". */
void hts_now_iso8601(char out[32]);
struct tm *convert_time_rfc822(struct tm *buffer, const char *s);
int set_filetime(const char *file, struct tm *tm_time);
int set_filetime_rfc822(const char *file, const char *date);

View File

@@ -547,6 +547,10 @@ struct httrackp {
archive. Tail: ABI */
hts_boolean warc_wacz; /**< --wacz: package archive+index+pages as a WACZ zip
(implies --warc + --warc-cdx). Tail: ABI */
hts_boolean changes; /**< --changes: report what this crawl changed against
the previous mirror. Tail: ABI */
void *changes_state; /**< live change-report accumulator (hts_changes*),
engine-owned. Tail: ABI */
};
/* Running statistics for a mirror. */

View File

@@ -460,6 +460,25 @@ static HTS_INLINE HTS_UNUSED const char *htsbuff_str(const htsbuff *b) {
return b->buf;
}
/**
* Copy src into dest (capacity size, NUL included), truncating to fit and
* always NUL-terminating. Unlike strlcpybuff() it never aborts, so it suits a
* value read back from a cache, a header or the wire, where refusing the whole
* record is worse than keeping a clipped one. Returns HTS_TRUE if it all fit;
* callers that clip on purpose ignore that, so it is not HTS_CHECK_RESULT.
*/
static HTS_INLINE HTS_UNUSED hts_boolean strclipbuff(char *dest, size_t size,
const char *src) {
size_t len, copy;
assertf(dest != NULL && src != NULL && size != 0);
len = strlen(src);
copy = len < size ? len : size - 1;
memcpy(dest, src, copy);
dest[copy] = '\0';
return copy == len ? HTS_TRUE : HTS_FALSE;
}
/**
* Callers that deliberately ignore truncation use this instead of
* slprintfbuff(), so it is not HTS_CHECK_RESULT.
@@ -469,6 +488,9 @@ static HTS_INLINE HTS_UNUSED HTS_PRINTF_FUN(3, 0) hts_boolean
int ret;
assertf(dest != NULL && size != 0);
/* a vsnprintf failing outright may write nothing at all, leaving whatever
the caller had on the stack for it to publish */
dest[0] = '\0';
ret = vsnprintf(dest, size, fmt, args);
/* pre-C99 runtimes (msvcrt _vsnprintf) return -1 and do not terminate */
dest[size - 1] = '\0';
@@ -492,6 +514,20 @@ static HTS_INLINE HTS_UNUSED HTS_CHECK_RESULT HTS_PRINTF_FUN(3, 4) hts_boolean
return ret;
}
/**
* slprintfbuff() for diagnostics quoting remote or client text, which are
* meant to be clipped: nothing to act on, hence not HTS_CHECK_RESULT. A (void)
* cast on slprintfbuff() is no substitute, GCC warns through it.
*/
static HTS_INLINE HTS_UNUSED HTS_PRINTF_FUN(3, 4) void slprintfbuff_clip(
char *dest, size_t size, const char *fmt, ...) {
va_list args;
va_start(args, fmt);
(void) vslprintfbuff(dest, size, fmt, args);
va_end(args);
}
/**
* slprintfbuff() over the in-scope array ARR (capacity = sizeof(ARR)).
* On GCC/Clang a pointer is a compile error; use slprintfbuff() for those.

View File

@@ -60,6 +60,7 @@ Please visit our Website: http://www.httrack.com
#include "htscodec.h"
#include "htsproxy.h"
#include "htswarc.h"
#include "htschanges.h"
#if HTS_USEZLIB
#include "htszlib.h"
#endif
@@ -457,6 +458,12 @@ static void basic_selftests(void) {
// link one level up -> a "../" prefix
assertf(lienrelatif(s, sizeof(s), "a.html", "dir/index.html") == 0);
assertf(strcmp(s, "../a.html") == 0);
// an empty current path: the trim used to walk off the front of it, which
// "?x" reaches too because the query pre-pass hands on the part before it
assertf(lienrelatif(s, sizeof(s), "dir/page.html", "") == 0);
assertf(strcmp(s, "dir/page.html") == 0);
assertf(lienrelatif(s, sizeof(s), "dir/page.html", "?x") == 0);
assertf(strcmp(s, "dir/page.html") == 0);
}
}
@@ -618,6 +625,66 @@ static int string_safety_selftests(void) {
return 1;
}
/* strclipbuff: truncate-and-report, never abort. Same canary shape; the
destination is poisoned before every call so a case cannot pass on the
previous one's bytes. */
{
struct {
char dst[8];
char canary[8];
} s;
memset(&s, '#', sizeof(s));
#define POISON_DST() memset(s.dst, '#', sizeof(s.dst))
/* well under capacity: no padding, nothing eaten off the end */
POISON_DST();
if (!strclipbuff(s.dst, sizeof(s.dst), "abc") || strcmp(s.dst, "abc") != 0)
return 1;
/* exact fit: capacity - 1 characters plus the NUL */
POISON_DST();
if (!strclipbuff(s.dst, sizeof(s.dst), "1234567") ||
strcmp(s.dst, "1234567") != 0)
return 1;
/* one over, then far over: clipped, terminated, reported. The expected
bytes differ from the case above, so a write-nothing implementation
cannot pass on the leftovers. */
POISON_DST();
if (strclipbuff(s.dst, sizeof(s.dst), "abcdefgh") ||
strcmp(s.dst, "abcdefg") != 0)
return 1;
POISON_DST();
if (strclipbuff(s.dst, sizeof(s.dst), "0123456789abcdef") ||
strcmp(s.dst, "0123456") != 0)
return 1;
/* degenerate capacity 1: only the NUL fits. Capacity 2 pins the boundary
between that and the sizes above: one character plus the NUL. */
POISON_DST();
if (strclipbuff(s.dst, 1, "x") || s.dst[0] != '\0')
return 1;
POISON_DST();
if (strclipbuff(s.dst, 2, "yz") || strcmp(s.dst, "y") != 0)
return 1;
/* the empty string fits any non-zero capacity */
POISON_DST();
if (!strclipbuff(s.dst, sizeof(s.dst), "") || s.dst[0] != '\0')
return 1;
/* a byte over 0x7f must not end the copy early */
POISON_DST();
if (!strclipbuff(s.dst, sizeof(s.dst), "\xff\xfe") ||
strcmp(s.dst, "\xff\xfe") != 0)
return 1;
#undef POISON_DST
if (memcmp(s.canary, "########", sizeof(s.canary)) != 0)
return 1;
}
/* htsblk_failf: clips a reason quoted from a remote reply into msg[] and
touches nothing else in the block */
{
@@ -1519,7 +1586,7 @@ static int st_hashtable(httrackp *opt, int argc, char **argv) {
size_t i;
for (i = bench[loop].offset; i < (size_t) count;
i += bench[loop].modulus) {
int result;
int result = 0; /* no final else: an unknown type reports failure */
FMT();
if (bench[loop].type == DO_ADD || bench[loop].type == DO_DRY_ADD) {
size_t k;
@@ -1674,6 +1741,12 @@ static int st_copyopt(httrackp *opt, int argc, char **argv) {
if (strcmp(StringBuff(to->warc_file), "run.warc.gz") != 0)
err = 1;
from->changes = HTS_TRUE;
to->changes = HTS_FALSE;
copy_htsopt(from, to);
if (to->changes != HTS_TRUE)
err = 1;
/* #185 pause pair: copied when enabled (max>0), the 0 sentinel skips */
from->pause_min_ms = 5000;
from->pause_max_ms = 10000;
@@ -5160,6 +5233,192 @@ static int st_cookieimport(httrackp *opt, int argc, char **argv) {
return 0;
}
/* --changes bucket accounting and JSON escaping (#714). */
static int st_changes(httrackp *opt, int argc, char **argv) {
String out = STRING_EMPTY;
int err = 0;
(void) opt;
(void) argc;
(void) argv;
/* A file the crawl did not rewrite is unchanged whatever the wire said. */
assertf(hts_changes_classify(HTS_FALSE, HTS_TRUE, HTS_FALSE, HTS_FALSE,
HTS_FALSE) == HTS_CHANGE_UNCHANGED);
/* Rewritten with no previous copy: new, digests or not. */
assertf(hts_changes_classify(HTS_TRUE, HTS_FALSE, HTS_FALSE, HTS_TRUE,
HTS_FALSE) == HTS_CHANGE_NEW);
/* Digests decide, and outrank the transfer signal both ways: a server with
no validators answers 200 with the same bytes, and a 304 can still sit in
front of a locally damaged copy. */
assertf(hts_changes_classify(HTS_TRUE, HTS_TRUE, HTS_FALSE, HTS_TRUE,
HTS_TRUE) == HTS_CHANGE_UNCHANGED);
assertf(hts_changes_classify(HTS_TRUE, HTS_TRUE, HTS_TRUE, HTS_TRUE,
HTS_FALSE) == HTS_CHANGE_CHANGED);
/* Only with no digest at all does the transfer signal get a say. */
assertf(hts_changes_classify(HTS_TRUE, HTS_TRUE, HTS_TRUE, HTS_FALSE,
HTS_FALSE) == HTS_CHANGE_UNCHANGED);
assertf(hts_changes_classify(HTS_TRUE, HTS_TRUE, HTS_FALSE, HTS_FALSE,
HTS_FALSE) == HTS_CHANGE_CHANGED);
#define JSON_IS(SRC, WANT) \
do { \
StringClear(out); \
hts_changes_json_string(&out, SRC); \
if (strcmp(StringBuff(out), WANT) != 0) { \
fprintf(stderr, "changes: %s -> %s, expected %s\n", #SRC, \
StringBuff(out), WANT); \
err = 1; \
} \
} while (0)
JSON_IS("/a/b.html", "\"/a/b.html\"");
JSON_IS("a\"b\\c", "\"a\\\"b\\\\c\"");
JSON_IS("tab\there", "\"tab\\u0009here\"");
/* Valid UTF-8 rides through; a lone Latin-1 byte, a truncated sequence and
an overlong encoding of '/' each become U+FFFD rather than invalid JSON. */
JSON_IS("caf\xc3\xa9", "\"caf\xc3\xa9\"");
JSON_IS("caf\xe9", "\"caf\\ufffd\"");
JSON_IS("\xc3", "\"\\ufffd\"");
JSON_IS("\xc0\xaf", "\"\\ufffd\\ufffd\"");
/* A UTF-16 surrogate half is well-formed UTF-8 by shape only. */
JSON_IS("\xed\xa0\x80", "\"\\ufffd\\ufffd\\ufffd\"");
#undef JSON_IS
StringFree(out);
printf("changes self-test: %s\n", err ? "FAIL" : "OK");
return err;
}
#define CHANGES_RACE_FILES 8
#define CHANGES_RACE_ROUNDS 400
static void changes_race_notify(httrackp *opt, int n) {
char fil[64];
char BIGSTK save[HTS_URLMAXSIZE * 2];
snprintf(fil, sizeof(fil), "/f%d.bin", n);
strlcpybuff(save, StringBuff(opt->path_html), sizeof(save));
strlcatbuff(save, "race.example", sizeof(save));
strlcatbuff(save, fil, sizeof(save));
hts_changes_notify(opt, "race.example", fil, save, HTS_TRUE, HTS_FALSE);
}
/* htsthread_wait() counts a thread only once it is running, so it can return
before any of them started; join on our own counter instead. */
static htsmutex changes_race_lock = HTSMUTEX_INIT;
static int changes_race_started = 0;
static int changes_race_live = 0;
static int changes_race_count(int *which) {
int n;
hts_mutexlock(&changes_race_lock);
n = *which;
hts_mutexrelease(&changes_race_lock);
return n;
}
static void changes_race_thread(void *arg) {
httrackp *const opt = (httrackp *) arg;
int i;
hts_mutexlock(&changes_race_lock);
changes_race_started++;
hts_mutexrelease(&changes_race_lock);
for (i = 0; i < CHANGES_RACE_ROUNDS; i++)
changes_race_notify(opt, i % CHANGES_RACE_FILES);
hts_mutexlock(&changes_race_lock);
changes_race_live--;
hts_mutexrelease(&changes_race_lock);
}
/* A transfer thread the crawl never joins (FTP) reaches hts_changes_notify()
while the report is being resolved and written. Run it under TSan. */
static int st_changes_race(httrackp *opt, int argc, char **argv) {
String out = STRING_EMPTY;
char base[HTS_URLMAXSIZE];
int err = 0;
int i;
if (argc < 1) {
fprintf(stderr, "usage: -#test=changes-race <writable directory>\n");
return 1;
}
strcpybuff(base, argv[0]);
if (base[0] != '\0' && base[strlen(base) - 1] != '/')
strcatbuff(base, "/");
StringCopy(opt->path_html, base);
StringCopy(opt->path_html_utf8, base);
StringCopy(opt->path_log, base);
opt->changes = HTS_TRUE;
hts_changes_free_opt(opt);
/* Real files, so the reader hashes and stats them for as long as it takes. */
{
char BIGSTK dir[HTS_URLMAXSIZE * 2];
strlcpybuff(dir, base, sizeof(dir));
strlcatbuff(dir, "race.example", sizeof(dir));
for (i = 0; i < CHANGES_RACE_FILES; i++) {
char BIGSTK path[HTS_URLMAXSIZE * 2];
char name[64];
FILE *fp;
int n;
snprintf(name, sizeof(name), "/f%d.bin", i);
strlcpybuff(path, dir, sizeof(path));
strlcatbuff(path, name, sizeof(path));
structcheck(path);
fp = FOPEN(path, "wb");
if (fp == NULL) {
fprintf(stderr, "changes-race: cannot write %s\n", path);
return 1;
}
for (n = 0; n < 16384; n++)
fwrite("0123456789abcdef", 1, 16, fp);
fclose(fp);
}
}
/* Take both locks once here: hts_mutexlock() initializes lazily, and two
threads reaching a fresh one together race on the init itself. */
hts_mutexlock(&changes_race_lock);
changes_race_started = 0;
changes_race_live = 4;
hts_mutexrelease(&changes_race_lock);
for (i = 0; i < 4; i++) {
if (hts_newthread(changes_race_thread, opt) != 0) {
fprintf(stderr, "changes-race: cannot spawn a notifier thread\n");
return 1;
}
}
/* Report only once they are all notifying, or there is nothing to race. */
while (changes_race_count(&changes_race_started) < 4)
Sleep(10);
for (i = 0; i < 64; i++)
hts_changes_report(opt, &out);
hts_changes_close_opt(opt);
while (changes_race_count(&changes_race_live) > 0)
Sleep(10);
/* Sealed: a straggler must be dropped, not start a report nobody writes. */
changes_race_notify(opt, CHANGES_RACE_FILES + 1);
hts_changes_report(opt, &out);
if (StringLength(out) == 0) {
fprintf(stderr, "changes-race: the report was lost after close\n");
err = 1;
} else if (strstr(StringBuff(out), "f9.bin") != NULL) {
fprintf(stderr, "changes-race: a post-close notify reached the report\n");
err = 1;
}
StringFree(out);
hts_changes_free_opt(opt);
printf("changes-race self-test: %s\n", err ? "FAIL" : "OK");
return err;
}
/* ------------------------------------------------------------ */
/* Registry: name -> handler, with a usage hint and a one-line description. */
/* ------------------------------------------------------------ */
@@ -5215,6 +5474,10 @@ static const struct selftest_entry {
{"strsafe", "[overflow|overflow-buff|overflow-src [str]]",
"bounded string-op self-test", st_strsafe},
{"copyopt", "", "copy_htsopt option-copy self-test", st_copyopt},
{"changes", "", "--changes bucket accounting and JSON escaping (#714)",
st_changes},
{"changes-race", "<dir>", "--changes under a late transfer thread (#714)",
st_changes_race},
{"pause", "", "randomized inter-file pause target self-test", st_pause},
{"relative", "<link> <curr-file>", "relative link between two paths",
st_relative},

View File

@@ -318,6 +318,18 @@ typedef struct {
error_redirect = "/server/error.html"; \
} while(0)
/* Longest error message shown on the error page; the rest is clipped. */
#define ERROR_MESSAGE_MAX 1024
/* SET_ERROR() with a printf format. Clips: these messages quote posted fields,
whose length the client picks. */
#define SET_ERRORF(...) \
do { \
char errbuf[ERROR_MESSAGE_MAX]; \
slprintfbuff_clip(errbuf, sizeof(errbuf), __VA_ARGS__); \
SET_ERROR(errbuf); \
} while (0)
/* Longest "sid" value worth unescaping: the expected one is an md5 hex digest,
so anything near this is already invalid and is rejected unread. */
#define SID_VALUE_MAX 64
@@ -725,7 +737,7 @@ int smallserver(T_SOC soc, char *url, char *method, char *data, char *path) {
fp = fopen(StringBuff(fspath), "rb");
if (fp) {
/* Read file */
while(!feof(fp)) {
while (!feof(fp) && !ferror(fp)) {
char *str = line;
char *pos;
@@ -879,28 +891,18 @@ int smallserver(T_SOC soc, char *url, char *method, char *data, char *path) {
commandEnd = 1;
}
} else {
char tmp[1024];
sprintf(tmp,
"Unable to write %d bytes in the the init file %s",
count, StringBuff(fspath));
SET_ERROR(tmp);
SET_ERRORF(
"Unable to write %d bytes in the the init file %s",
count, StringBuff(fspath));
}
fclose(fp);
} else {
char tmp[1024];
sprintf(tmp, "Unable to create the init file %s",
StringBuff(fspath));
SET_ERROR(tmp);
SET_ERRORF("Unable to create the init file %s",
StringBuff(fspath));
}
} else {
char tmp[1024];
sprintf(tmp,
"Unable to create the directory structure in %s",
StringBuff(fspath));
SET_ERROR(tmp);
SET_ERRORF("Unable to create the directory structure in %s",
StringBuff(fspath));
}
} else {
@@ -978,10 +980,10 @@ int smallserver(T_SOC soc, char *url, char *method, char *data, char *path) {
}
}
/* path itself may hold ".." (webhttrack passes "<bin>/../share"), so
only the untrusted halves are checked: file here, website above. */
if (fsfile[0] && strstr(file, "..") == NULL
&& (fp = fopen(fsfile, "rb"))) {
/* Regular files only: reading a directory or FIFO never ends, and
"path" may hold "..", so only the untrusted halves are checked. */
if (fsfile[0] && strstr(file, "..") == NULL && fexist(fsfile) &&
(fp = fopen(fsfile, "rb"))) {
char ok[] =
"HTTP/1.0 200 OK\r\n" "Connection: close\r\n"
"Server: httrack-small-server\r\n" "Content-type: text/html\r\n"
@@ -1040,7 +1042,7 @@ int smallserver(T_SOC soc, char *url, char *method, char *data, char *path) {
int outputmode = 0;
StringMemcat(headers, ok, sizeof(ok) - 1);
while(!feof(fp)) {
while (!feof(fp) && !ferror(fp)) {
char *str = line;
int prevlen = (int) StringLength(output);
int nocr = 0;
@@ -1474,14 +1476,15 @@ int smallserver(T_SOC soc, char *url, char *method, char *data, char *path) {
while(!feof(fp)) {
int n = (int) fread(line, 1, sizeof(line) - 2, fp);
if (n > 0) {
StringMemcat(output, line, n);
if (n <= 0) {
break; /* short read: EOF or error, never a retry */
}
StringMemcat(output, line, n);
}
}
fclose(fp);
} else if (strcmp(file, "/ping") == 0
|| strncmp(file, "/ping?", 6) == 0) {
} else if (strcmp(file, "/ping") == 0 ||
strncmp(file, "/ping?", 6) == 0) {
char error_hdr[] =
"HTTP/1.0 200 Pong\r\n" "Server: httrack small server\r\n"
"Content-type: text/html\r\n";

63
src/htsstats.h Normal file
View File

@@ -0,0 +1,63 @@
/* ------------------------------------------------------------ */
/*
HTTrack Website Copier, Offline Browser for Windows and Unix
Copyright (C) 1998 Xavier Roche and other contributors
SPDX-License-Identifier: GPL-3.0-or-later
This program is free software: you can redistribute it and/or modify
it under the terms of the GNU General Public License as published by
the Free Software Foundation, either version 3 of the License, or
(at your option) any later version.
This program is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
GNU General Public License for more details.
You should have received a copy of the GNU General Public License
along with this program. If not, see <http://www.gnu.org/licenses/>.
Ethical use: we kindly ask that you NOT use this software to harvest email
addresses or to collect any other private information about people. Doing so
would dishonor our work and waste the many hours we have spent on it.
Please visit our Website: http://www.httrack.com
*/
/* ------------------------------------------------------------ */
/* File: in-progress display rows, shared by httrack and htsserver .h */
/* Author: Xavier Roche */
/* ------------------------------------------------------------ */
#ifndef HTS_STATS_DEFH
#define HTS_STATS_DEFH
#include "htsglobal.h"
/* One row of the "in progress" display, shared so the CLI (httrack) and the
web GUI (htsserver) cannot drift apart: each fills its own array. */
#define NStatsBuffer 14
#ifndef HTS_DEF_FWSTRUCT_t_StatsBuffer
#define HTS_DEF_FWSTRUCT_t_StatsBuffer
typedef struct t_StatsBuffer t_StatsBuffer;
#endif
struct t_StatsBuffer {
char name[1024];
char file[1024];
char state[288]; // a short label plus back->info[256]
char BIGSTK url_sav[HTS_URLMAXSIZE * 2]; // pour cancel
char BIGSTK url_adr[HTS_URLMAXSIZE * 2];
char BIGSTK url_fil[HTS_URLMAXSIZE * 2];
LLint size;
LLint sizetot;
int offset;
//
int back;
//
int actived; // pour disabled
};
#endif

View File

@@ -308,8 +308,10 @@ int lienrelatif(char *s, size_t ssize, const char *link, const char *curr_fil) {
// copy only the current path
curr = _curr;
strlcpybuff(curr, curr_fil, sizeof(_curr));
if ((a = strchr(curr, '?')) == NULL) // couper au ? (params)
a = curr + strlen(curr) - 1; // pas de params: aller à la fin
if ((a = strchr(curr, '?')) == NULL) { // cut at the ? (query parameters)
// an empty path has no last character: curr-1 would read before the buffer
a = curr[0] != '\0' ? curr + strlen(curr) - 1 : curr;
}
while((*a != '/') && (a > curr))
a--; // chercher dernier / du chemin courant
if (*a == '/')

View File

@@ -462,22 +462,6 @@ static void warc_make_id(warc_writer *w, char out[64]) {
b[11], b[12], b[13], b[14], b[15]);
}
static void warc_now_iso8601(char out[32]) {
time_t t = time(NULL);
struct tm tmv;
#if defined(_WIN32)
struct tm *g = gmtime(&t);
if (g != NULL)
tmv = *g;
else
memset(&tmv, 0, sizeof(tmv));
#else
if (gmtime_r(&t, &tmv) == NULL)
memset(&tmv, 0, sizeof(tmv));
#endif
strftime(out, 32, "%Y-%m-%dT%H:%M:%SZ", &tmv);
}
/* Case-insensitive "is this the header named name?" test, tolerating optional
whitespace before the ':' (non-compliant "Name : value" is still matched). */
static int header_is(const char *line, size_t line_len, const char *name) {
@@ -1063,7 +1047,7 @@ static void warc_wacz_package(warc_writer *w) {
wbuf_free(&pages);
/* datapackage.json listing every stored file with its sha256 + size. */
warc_now_iso8601(created);
hts_now_iso8601(created);
memset(&dp, 0, sizeof(dp));
if (!err &&
(wbuf_printf(&dp,
@@ -1166,7 +1150,7 @@ static int warc_emit(warc_writer *w, const char *type, const char *content_type,
memset(&hdr, 0, sizeof(hdr));
warc_make_id(w, id);
warc_now_iso8601(date);
hts_now_iso8601(date);
#if HTS_USEOPENSSL
/* Block digest over the whole block, in one streaming pass. */
@@ -1670,7 +1654,7 @@ int warc_write_transaction(warc_writer *w, const char *target_uri,
char date[32];
char *slot;
size_t need;
warc_now_iso8601(date);
hts_now_iso8601(date);
need = strlen(target_uri) + 1 + strlen(date) + 1;
slot = malloct(need);
if (slot != NULL) {

View File

@@ -35,26 +35,10 @@ Please visit our Website: http://www.httrack.com
#include "htsglobal.h"
#include "htscore.h"
#include "htsstats.h"
#define NStatsBuffer 14
#define MAX_LEN_INPROGRESS 40
typedef struct t_StatsBuffer {
char name[1024];
char file[1024];
char state[288]; // a short label plus back->info[256]
char url_sav[HTS_URLMAXSIZE * 2]; // pour cancel
char url_adr[HTS_URLMAXSIZE * 2];
char url_fil[HTS_URLMAXSIZE * 2];
LLint size;
LLint sizetot;
int offset;
//
int back;
//
int actived; // pour disabled
} t_StatsBuffer;
typedef struct t_InpInfo {
int ask_refresh;
int refresh;

View File

@@ -230,8 +230,7 @@ static hts_boolean vt_size_refresh(void) {
*/
#define STYLE_STATVALUES VT_BOLD
#define STYLE_STATTEXT VT_UNBOLD
#define STYLE_STATRESET VT_UNBOLD
#define NStatsBuffer 14
#define STYLE_STATRESET VT_UNBOLD
/* Rows the stats block and "Current job" take above the in-progress list. */
#define NStatsHeaderRows 7

View File

@@ -36,26 +36,7 @@ Please visit our Website: http://www.httrack.com
#include "htsglobal.h"
#include "htscore.h"
#include "htssafe.h"
#ifndef HTS_DEF_FWSTRUCT_t_StatsBuffer
#define HTS_DEF_FWSTRUCT_t_StatsBuffer
typedef struct t_StatsBuffer t_StatsBuffer;
#endif
struct t_StatsBuffer {
char name[1024];
char file[1024];
char state[288]; // a short label plus back->info[256]
char BIGSTK url_sav[HTS_URLMAXSIZE * 2]; // pour cancel
char BIGSTK url_adr[HTS_URLMAXSIZE * 2];
char BIGSTK url_fil[HTS_URLMAXSIZE * 2];
LLint size;
LLint sizetot;
int offset;
//
int back;
//
int actived; // pour disabled
};
#include "htsstats.h"
#ifndef HTS_DEF_FWSTRUCT_t_InpInfo
#define HTS_DEF_FWSTRUCT_t_InpInfo

View File

@@ -138,6 +138,7 @@
<ClCompile Include="htswrap.c" />
<ClCompile Include="htszlib.c" />
<ClCompile Include="htswarc.c" />
<ClCompile Include="htschanges.c" />
<ClCompile Include="md5.c" />
<ClCompile Include="minizip\ioapi.c" />
<ClCompile Include="minizip\iowin32.c" />

View File

@@ -820,8 +820,8 @@ static PT_Element proxytrack_process_DAV_Request(PT_Indexes indexes,
elt->size = StringLength(response);
elt->adr = StringAcquire(&response);
elt->statuscode = 207; /* Multi-Status */
strcpy(elt->charset, "utf-8");
strcpy(elt->contenttype, "text/xml");
strcpybuff(elt->charset, "utf-8");
strcpybuff(elt->contenttype, "text/xml");
strcpybuff(elt->msg, "Multi-Status");
StringFree(response);
@@ -877,8 +877,8 @@ static PT_Element proxytrack_process_HTTP_List(PT_Indexes indexes,
elt->size = StringLength(html);
elt->adr = StringAcquire(&html);
elt->statuscode = HTTP_OK;
strcpy(elt->charset, "iso-8859-1");
strcpy(elt->contenttype, "text/html");
strcpybuff(elt->charset, "iso-8859-1");
strcpybuff(elt->contenttype, "text/html");
strcpybuff(elt->msg, "OK");
StringFree(html);
return elt;

View File

@@ -368,7 +368,11 @@ HTS_UNUSED static struct tm PT_GetTime(time_t t) {
if (tm != NULL)
return *tm;
else {
/* an all-zero tm has tm_mday == 0, which the ARC date field prints as a
day of "00"; the epoch is the conventional "date unknown" */
memset(&tmbuf, 0, sizeof(tmbuf));
tmbuf.tm_year = 70;
tmbuf.tm_mday = 1;
return tmbuf;
}
}

View File

@@ -405,8 +405,8 @@ PT_Element PT_Index_HTML_BuildRootInfo(PT_Indexes indexes) {
elt->size = StringLength(html);
elt->adr = StringAcquire(&html);
elt->statuscode = HTTP_OK;
strcpy(elt->charset, "iso-8859-1");
strcpy(elt->contenttype, "text/html");
strcpybuff(elt->charset, "iso-8859-1");
strcpybuff(elt->contenttype, "text/html");
strcpybuff(elt->msg, "OK");
StringFree(html);
return elt;
@@ -829,16 +829,8 @@ PT_Element PT_ElementNew(void) {
}
/* ProxyTrack's htsblk_failf(): a clipped, diagnostic-only failure reason. */
static void PT_Element_failf(PT_Element r, const char *fmt, ...)
HTS_PRINTF_FUN(2, 3);
static void PT_Element_failf(PT_Element r, const char *fmt, ...) {
va_list args;
va_start(args, fmt);
(void) vslprintfbuff(r->msg, sizeof(r->msg), fmt, args);
va_end(args);
}
#define PT_Element_failf(R, ...) \
slprintfbuff_clip((R)->msg, sizeof((R)->msg), __VA_ARGS__)
PT_Element PT_ReadCache(PT_Index index, const char *url, int flags) {
if (index != NULL && SAFE_INDEX(index)) {
@@ -879,13 +871,11 @@ static PT_Element PT_ReadCache__New(PT_Index index, const char *url, int flags)
(headersSize) += (int) strlen(headers + headersSize); \
} while(0)
/* refvalue_size is mandatory: the cache line is bounded only by the line
buffer, not by the destination. Clip rather than reject, since an
engine-written field can be wider than ours. */
buffer, not by the destination. */
#define ZIP_READFIELD_STRING(line, value, refline, refvalue, refvalue_size) \
do { \
if (line[0] != '\0' && strfield2(line, refline)) { \
(refvalue)[0] = '\0'; \
strlncatbuff(refvalue, value, refvalue_size, (refvalue_size) - 1); \
(void) strclipbuff(refvalue, refvalue_size, value); \
line[0] = '\0'; \
} \
} while (0)
@@ -1042,7 +1032,7 @@ static PT_Element PT_ReadCache__New_u(PT_Index index_, const char *url,
previous_save[0] = previous_save_[0] = '\0';
memset(r, 0, sizeof(_PT_Element));
r->location = location_default;
strcpy(r->location, "");
r->location[0] = '\0';
if (strncmp(url, "http://", 7) == 0)
url += 7;
hash_pos_return = coucal_read(index->hash, url, &hash_pos);
@@ -1127,7 +1117,7 @@ static PT_Element PT_ReadCache__New_u(PT_Index index_, const char *url,
int pathLen = (int) strlen(index->path);
if (pathLen > 0 && strncmp(previous_save_, index->path, pathLen) == 0) { // old (<3.40) buggy format
strcpy(previous_save, previous_save_);
strcpybuff(previous_save, previous_save_);
}
// relative ? (hack)
else if (index->safeCache || (previous_save_[0] != '/' // /home/foo/bar.gif
@@ -1554,7 +1544,7 @@ static int PT_LoadCache__Old(PT_Index index_, const char *filename) {
cache->version = (int) (firstline[8] - '0'); // cache 1.x
if (cache->version <= 5) {
a += cache_brstr(a, firstline, sizeof(firstline));
strcpy(cache->lastmodified, firstline);
strcpybuff(cache->lastmodified, firstline);
} else {
fclose(cache->dat);
cache->dat = NULL;
@@ -1571,7 +1561,7 @@ static int PT_LoadCache__Old(PT_Index index_, const char *filename) {
} else { // Vieille version du cache
/* */
cache->version = 0; // cache 1.0
strcpy(cache->lastmodified, firstline);
strcpybuff(cache->lastmodified, firstline);
}
/* Create hash table for the cache (MUCH FASTER!) */
@@ -1587,7 +1577,14 @@ static int PT_LoadCache__Old(PT_Index index_, const char *filename) {
a++;
/* read "host/file" */
a += binput(a, line, HTS_URLMAXSIZE);
a += binput(a, line + strlen(line), HTS_URLMAXSIZE);
{
/* binput writes its NUL at s[max], so the second field must be
bounded by what is left of line[], not by the same constant
*/
const size_t used = strlen(line);
a += binput(a, line + used, (int) (sizeof(line) - used - 1));
}
/* read position */
a += binput(a, linepos, 200);
sscanf(linepos, "%d", &pos);
@@ -1684,7 +1681,7 @@ static PT_Element PT_ReadCache__Old_u(PT_Index index_, const char *url,
previous_save[0] = previous_save_[0] = '\0';
memset(r, 0, sizeof(_PT_Element));
r->location = location_default;
strcpy(r->location, "");
r->location[0] = '\0';
if (strncmp(url, "http://", 7) == 0)
url += 7;
hash_pos_return = coucal_read(cache->hash, url, &hash_pos);
@@ -1791,7 +1788,7 @@ static PT_Element PT_ReadCache__Old_u(PT_Index index_, const char *url,
int pathLen = (int) strlen(index->path);
if (pathLen > 0 && strncmp(previous_save_, index->path, pathLen) == 0) { // old (<3.40) buggy format
strcpy(previous_save, previous_save_);
strcpybuff(previous_save, previous_save_);
}
// relative ? (hack)
else if (index->safeCache || (previous_save_[0] != '/' // /home/foo/bar.gif
@@ -2168,8 +2165,7 @@ int PT_LoadCache__Arc(PT_Index index_, const char *filename) {
#define HTTP_READFIELD_STRING(line, value, refline, refvalue, refvalue_size) \
do { \
if (line[0] != '\0' && strfield2(line, refline)) { \
(refvalue)[0] = '\0'; \
strlncatbuff(refvalue, value, refvalue_size, (refvalue_size) - 1); \
(void) strclipbuff(refvalue, refvalue_size, value); \
line[0] = '\0'; \
} \
} while (0)
@@ -2208,7 +2204,7 @@ static PT_Element PT_ReadCache__Arc_u(PT_Index index_, const char *url,
location_default[0] = '\0';
memset(r, 0, sizeof(_PT_Element));
r->location = location_default;
strcpy(r->location, "");
r->location[0] = '\0';
if (strncmp(url, "http://", 7) == 0)
url += 7;
hash_pos_return = coucal_read(index->hash, url, &hash_pos);
@@ -2380,8 +2376,18 @@ static int PT_SaveCache__Arc_Fun(void *arg, const char *url, PT_Element element)
PT_SaveCache__Arc_t *st = (PT_SaveCache__Arc_t *) arg;
FILE *const fp = st->fp;
struct tm *tm = convert_time_rfc822(&st->buff, element->lastmodified);
struct tm unknown_date;
int size_headers;
/* a cached entry with no parseable Last-Modified must not take the writer
down; the epoch is the conventional "date unknown" */
if (tm == NULL) {
memset(&unknown_date, 0, sizeof(unknown_date));
unknown_date.tm_year = 70;
unknown_date.tm_mday = 1;
tm = &unknown_date;
}
sprintf(st->headers,
"HTTP/1.0 %d %s"
"\r\n"
@@ -2417,7 +2423,7 @@ static int PT_SaveCache__Arc_Fun(void *arg, const char *url, PT_Element element)
if (element->adr != NULL) {
domd5mem(element->adr, element->size, st->md5, 1);
} else {
strcpy(st->md5, "-");
strcpybuff(st->md5, "-");
}
fprintf(fp,
/* nl */

View File

@@ -0,0 +1,30 @@
#!/bin/bash
#
set -euo pipefail
tmpdir=$(mktemp -d "${TMPDIR:-/tmp}/httrack_changes_st.XXXXXX") || exit 1
trap 'rm -rf "$tmpdir"' EXIT HUP INT QUIT PIPE TERM
# No pipe into grep: SIGPIPE would mask a failing exit status.
expect_ok() {
local label="$1" out
shift
out=$("$@" 2>&1) || {
echo "FAIL: ${label} exited non-zero: ${out}"
exit 1
}
case "$out" in
*"${label}: OK"*) ;;
*)
echo "FAIL: ${out}"
exit 1
;;
esac
}
# --changes bucket accounting and JSON escaping (#714).
expect_ok "changes self-test" httrack -O "${tmpdir}/o1" -#test=changes run
# The report must hold against a transfer thread the crawl never joins.
expect_ok "changes-race self-test" httrack -O "${tmpdir}/o2" \
-#test=changes-race "${tmpdir}/race"

View File

@@ -146,11 +146,13 @@ sid=$(scrape_sid "${port}")
resp=$(post_redirect '/foo
X-Injected: pwned' "${port}" "${sid}") || fail "no response to a CRLF redirect value"
stop
printf '%s' "${resp}" | grep -qi 'X-Injected' &&
# Here-string, not a pipe: grep -q SIGPIPEs the producer on the first match, and
# pipefail then turns the hit into a silent pass.
grep -qi 'X-Injected' <<<"${resp}" &&
fail "CRLF in the redirect value reached the response"
test "$(printf '%s' "${resp}" | grep -ci '^Location:')" -eq 0 ||
test "$(grep -ci '^Location:' <<<"${resp}")" -eq 0 ||
fail "CRLF redirect still emitted a Location header"
test "$(printf '%s' "${resp}" | grep -c '^HTTP/1\.')" -eq 1 ||
test "$(grep -c '^HTTP/1\.' <<<"${resp}")" -eq 1 ||
fail "CRLF redirect did not yield exactly one status line"
# Loopback unless asked otherwise. Assert the socket, not the announcement.

View File

@@ -103,9 +103,10 @@ sid=$(scrape_sid "${port}")
test "${#sid}" -eq 32 || fail "did not scrape a 32-hex sid from the page (got '${sid}')"
# Accept: the legitimate flow still works. "redirect" is the cheapest field
# with a reply that is visible in the headers.
# with a reply that is visible in the headers. Matched from a here-string, not a
# pipe: grep -q SIGPIPEs the producer, and pipefail then buries the verdict.
resp=$(request "${port}" "sid=${sid}&redirect=/accepted")
printf '%s' "${resp}" | grep -q '^Location: /accepted' ||
grep -q '^Location: /accepted' <<<"${resp}" ||
fail "a body carrying the correct sid was refused"
# Refuse: missing and wrong. A missing one is the regression under test — the
@@ -114,12 +115,12 @@ printf '%s' "${resp}" | grep -q '^Location: /accepted' ||
for bad in "redirect=/nosid" "sid=&redirect=/empty" \
"sid=00000000000000000000000000000000&redirect=/wrong"; do
resp=$(request "${port}" "${bad}")
printf '%s' "${resp}" | grep -q '^Location:' &&
grep -q '^Location:' <<<"${resp}" &&
fail "body accepted without a valid sid: ${bad}"
# A refusal has to be a well-formed reply, not a headerless fragment: the
# release build used to emit only Content-length, which reads as a protocol
# error to any client and hides the reason.
printf '%s' "${resp}" | grep -q '^HTTP/1\.0 403 ' ||
grep -q '^HTTP/1\.0 403 ' <<<"${resp}" ||
fail "refusal was not a 403: ${bad}"
done
@@ -127,16 +128,23 @@ done
# applied to one global key store that later requests render from, and the
# command dispatcher reads it from there. Probe the store itself — step3
# interpolates ${projname} into its title — rather than the refused reply.
title() { request "$1" "" step3; }
store_page() {
# a dead probe or a non-page reply must not read as "the store was not written"
page=$(request "$1" "" step3) || fail "the key store probe failed"
grep -q '^HTTP/1\.0 200 ' <<<"${page}" ||
fail "the key store probe did not answer: '$(head -1 <<<"${page}")'"
}
request "${port}" "projname=UNAUTHWRITE" >/dev/null 2>&1 || true
title "${port}" | grep -q 'UNAUTHWRITE' &&
store_page "${port}"
grep -q 'UNAUTHWRITE' <<<"${page}" &&
fail "a body without a valid sid was written to the key store"
# Paired accept case: without it the assertion above passes even if the server
# simply ignores every body, which would prove nothing.
request "${port}" "sid=${sid}&projname=AUTHWRITE" >/dev/null 2>&1 || true
title "${port}" | grep -q 'AUTHWRITE' ||
store_page "${port}"
grep -q 'AUTHWRITE' <<<"${page}" ||
fail "a body carrying the correct sid was not written to the key store"
echo "PASS"

View File

@@ -90,6 +90,15 @@ sys.stdout.write(out.decode("latin-1"))' "$1" "$2" "$3"
get() { request "$1" "$2" ""; }
post() { request "$1" "" "$2"; }
# GET $2 into ${reply}, requiring status $3. Captured, not piped: a dead request
# reads as marker-absent, and so do a truncated body and a 302 to the file.
reply=
fetch() {
reply=$(get "$1" "$2") || fail "the request for $2 failed"
grep -q "^HTTP/1\.0 $3 " <<<"${reply}" ||
fail "$2: wanted a $3 reply, got '$(head -1 <<<"${reply}")'"
}
url=$(start)
port=$(portof "${url}")
srv=$(srvpid)
@@ -109,13 +118,14 @@ echo "LOGMARKER" >"${base}/proj/hts-log.txt"
# with a component over NAME_MAX (mkdir refuses it whatever the uid). Must precede
# any successful save: commandEnd then swaps the error page for the finished one.
get "${port}" /server/style.css >/dev/null # leaves a path behind in fsfile
post "${port}" "sid=${sid}&command=httrack&command_do=save&winprofile=x&path=${base}/$(printf '%0300d' 0)&projname=p" |
grep -q '^Location: /server/error.html' ||
saved=$(post "${port}" "sid=${sid}&command=httrack&command_do=save&winprofile=x&path=${base}/$(printf '%0300d' 0)&projname=p")
grep -q '^Location: /server/error.html' <<<"${saved}" ||
fail "the refused save did not redirect to the error page"
# No project yet, so no root to serve from.
post "${port}" "sid=${sid}&projpath=${base}/" >/dev/null
get "${port}" /website/secret.txt | grep -q SECRETMARKER &&
fetch "${port}" /website/secret.txt 404
grep -q SECRETMARKER <<<"${reply}" &&
fail "a posted projpath served a file outside any project"
# Positive control: step4's "save settings" flow registers the project without
@@ -126,24 +136,29 @@ body="${body}&path=${base}&projname=proj&projpath=${base}/proj/"
post "${port}" "${body}" >/dev/null
test -f "${base}/proj/hts-cache/winprofile.ini" ||
fail "the project was not registered: $(cat "${log}")"
get "${port}" /website/hts-log.txt | grep -q LOGMARKER ||
fetch "${port}" /website/hts-log.txt 200
grep -q LOGMARKER <<<"${reply}" ||
fail "the registered project's mirror is not served"
# Same request with the root repointed: the project is legitimate, projpath is
# not what decides where the bytes come from.
post "${port}" "sid=${sid}&projpath=${base}/" >/dev/null
get "${port}" /website/secret.txt | grep -q SECRETMARKER &&
fetch "${port}" /website/secret.txt 404
grep -q SECRETMARKER <<<"${reply}" &&
fail "a posted projpath repointed the served root"
post "${port}" "sid=${sid}&projpath=/etc/" >/dev/null
get "${port}" /website/passwd | grep -q '^root:' &&
fetch "${port}" /website/passwd 404
grep -q '^root:' <<<"${reply}" &&
fail "a posted projpath read an arbitrary system file"
# A ".." anywhere in the recorded root would escape the mirror on every later
# request, so the save must be refused and the previous root kept.
post "${port}" "sid=${sid}&command=httrack&command_do=save&winprofile=x&path=${base}/proj&projname=.." >/dev/null
get "${port}" /website/secret.txt | grep -q SECRETMARKER &&
fetch "${port}" /website/secret.txt 404
grep -q SECRETMARKER <<<"${reply}" &&
fail "a '..' in the saved project path escaped the mirror root"
get "${port}" /website/hts-log.txt | grep -q LOGMARKER ||
fetch "${port}" /website/hts-log.txt 200
grep -q LOGMARKER <<<"${reply}" ||
fail "rejecting the '..' root also lost the previous one"
# The root is now the server's own, but it is still built from two posted
@@ -162,7 +177,7 @@ test -f "${fspath}/hts-cache/winprofile.ini" ||
fail "the long-path project was not registered: $(cat "${log}")"
get "${port}" "/website/${longurl}" >/dev/null 2>&1 || true
alive "${srv}" || fail "an over-long project path crashed the server: $(cat "${log}")"
get "${port}" /server/index.html | grep -q '200 OK' ||
fail "the server stopped answering after the over-long project path"
# Not just alive: still answering.
fetch "${port}" /server/index.html 200
echo "PASS"

View File

@@ -0,0 +1,51 @@
#!/bin/bash
# A cached entry whose Last-Modified is missing or unparseable must still
# convert: the ARC writer used to dereference the parsed date unchecked.
set -euo pipefail
dir=$(mktemp -d)
trap 'rm -rf "$dir"' EXIT
# $1 Last-Modified value (empty = omit the header), $2 expected archive date,
# $3 label. The date is asserted exactly: a guard that fires unconditionally
# would clobber the valid case, and one that skips the year/day would emit a
# month and day of 00.
run_case() {
lastmod=$1
wantdate=$2
what=$3
{
printf 'HTTP/1.1 200 OK\r\n'
printf 'Content-Type: text/html\r\n'
test -z "$lastmod" || printf 'Last-Modified: %s\r\n' "$lastmod"
printf 'Content-Length: 5\r\n\r\n'
} >"$dir/hdr"
printf 'hello' >"$dir/body"
alen=$(($(wc -c <"$dir/hdr") + $(wc -c <"$dir/body")))
{
printf 'filedesc://t.arc 0.0.0.0 20250101000000 text/plain 200 - - 0 t.arc 9\n'
printf '2 0 test\n'
printf '\n\n'
printf 'http://example.com/p.html 0.0.0.0 20250101000000 text/html 200 - - 0 t.arc %d\n' "$alen"
cat "$dir/hdr" "$dir/body"
} >"$dir/in.arc"
proxytrack --convert "$dir/out.arc" "$dir/in.arc" >/dev/null 2>&1 || {
echo "FAIL: proxytrack crashed on $what" >&2
exit 1
}
grep -aq "^http://example.com/p.html 0.0.0.0 ${wantdate} " "$dir/out.arc" || {
echo "FAIL: $what: expected archive date ${wantdate}, got:" >&2
grep -a '^http://example.com/' "$dir/out.arc" >&2 || echo "(entry dropped)" >&2
exit 1
}
}
# the epoch stands in for "date unknown"
run_case '' 19700101000000 'a missing Last-Modified'
run_case 'not a date at all' 19700101000000 'an unparseable Last-Modified'
run_case '0' 19700101000000 'a bare 0 Last-Modified'
# a parseable date must still come through untouched
run_case 'Sun, 06 Nov 1994 08:49:37 GMT' 19941106084937 'a valid Last-Modified'

View File

@@ -0,0 +1,69 @@
#!/bin/bash
# A cache file whose mtime is beyond what gmtime can represent must not put a
# day of "00" in the ARC filedesc date: PT_GetTime fell back to an all-zero tm.
set -euo pipefail
testdir=$(cd "$(dirname "$0")" && pwd)
# shellcheck source=tests/testlib.sh
. "${testdir}/testlib.sh"
python=$(find_python) || {
echo "python3 not found; skipping" >&2
exit 77
}
dir=$(mktemp -d)
trap 'rm -rf "$dir"' EXIT
rc=0
"$python" - "$dir/in.zip" <<'EOF' || rc=$?
import os, sys, zipfile
body = b"hello world"
meta = (
"X-In-Cache: 1\r\n"
"X-StatusCode: 200\r\n"
"X-StatusMessage: OK\r\n"
"X-Size: %d\r\n"
"Content-Type: text/html\r\n"
"Last-Modified: Wed, 01 Jan 2025 00:00:00 GMT\r\n"
"X-Save: out/page.html\r\n" % len(body)
).encode()
zi = zipfile.ZipInfo("example.com/page.html")
zi.compress_type = zipfile.ZIP_STORED
zi.extra = meta
with zipfile.ZipFile(sys.argv[1], "w") as z:
z.writestr(zi, body)
# past gmtime's range; the index timestamp is this file's mtime
huge = 4611686018427387903
try:
os.utime(sys.argv[1], (huge, huge))
except OSError:
sys.exit(77)
if int(os.stat(sys.argv[1]).st_mtime) < huge:
sys.exit(77) # filesystem clamped it, nothing to test
EOF
test "$rc" -eq 0 || {
test "$rc" -eq 77 || {
echo "FAIL: could not build the cache (python exit $rc)" >&2
exit 1
}
echo "filesystem will not hold an out-of-range mtime; skipping" >&2
exit 77
}
proxytrack --convert "$dir/out.arc" "$dir/in.zip" >/dev/null 2>&1 || {
echo "FAIL: proxytrack failed on a cache with an out-of-range mtime" >&2
exit 1
}
# field 3 of the filedesc line is YYYYMMDDhhmmss; the epoch stands in for
# "unknown", and 14 digits with a real day is the point (00 is what broke)
date=$(head -n1 "$dir/out.arc" | awk '{ print $3 }')
test "$date" = 19700101000000 || {
echo "FAIL: filedesc date is '$date', expected 19700101000000" >&2
exit 1
}

View File

@@ -0,0 +1,174 @@
#!/bin/bash
#
# The "save settings" failure messages quote a project path the request body
# supplies, so composing them must clip instead of running off a fixed buffer.
set -euo pipefail
testdir=$(cd "$(dirname "$0")" && pwd)
distdir=${top_srcdir:-$(cd "${testdir}/.." && pwd)}
distdir=$(cd "${distdir}" && pwd)
fail() {
echo "FAIL: $*" >&2
exit 1
}
command -v htsserver >/dev/null || fail "no htsserver in PATH"
command -v python3 >/dev/null || {
echo "python3 not found; skipping" >&2
exit 77
}
# ERROR_MESSAGE_MAX - 1 in htsserver.c.
msgmax=1023
srv=
log=$(mktemp)
# Physical path: macOS resolves TMPDIR through /var -> private/var, and the
# paths built below sit within a few bytes of its 1024-byte PATH_MAX.
base=$(cd "$(mktemp -d)" && pwd -P)
cleanup() {
test -z "${srv}" || kill -9 "${srv}" 2>/dev/null || true
rm -f "${log}"
rm -rf "${base}"
}
trap cleanup EXIT HUP INT QUIT PIPE TERM
freeport() {
python3 -c 'import socket
s = socket.socket()
s.bind(("127.0.0.1", 0))
print(s.getsockname()[1])
s.close()'
}
# Echo the announced URL. Runs in a command substitution, so it is a
# subshell and cannot export the pid: the caller reads it back with srvpid.
start() {
local port url
port=$(freeport)
: >"${log}"
(
trap '' TERM TTOU XFSZ
# Caps regular-file writes, which is how the failing-write branch below
# is reached without /dev/full; the announce rides a pipe, exempt.
ulimit -f 8
exec htsserver "${distdir}/" --port "${port}" 2>&1
) | cat >"${log}" &
for _ in $(seq 1 40); do
url=$(sed -n 's/^URL=//p' "${log}" 2>/dev/null) && test -n "${url}" && break
sleep 0.25
done
test -n "${url:-}" || fail "htsserver did not come up: $(cat "${log}")"
echo "${url}"
}
srvpid() { sed -n 's/^PID=//p' "${log}" | head -1; }
alive() { kill -0 "$1" 2>/dev/null; }
portof() { echo "${1##*:}" | tr -d /; }
# Raw request to 127.0.0.1:$1: GET the path $2, or POST the body $3 to / when
# $2 is empty. Prints the reply.
request() {
python3 -c 'import socket, sys
port, path, body = int(sys.argv[1]), sys.argv[2], sys.argv[3]
if path:
req = "GET %s HTTP/1.0\r\nHost: 127.0.0.1\r\n\r\n" % path
else:
req = ("POST / HTTP/1.0\r\nHost: 127.0.0.1\r\n"
"Content-type: application/x-www-form-urlencoded\r\n"
"Content-length: %d\r\n\r\n%s" % (len(body), body))
s = socket.create_connection(("127.0.0.1", port), 10)
s.settimeout(30)
s.sendall(req.encode())
out = b""
while True:
b = s.recv(65536)
if not b:
break
out += b
s.close()
sys.stdout.write(out.decode("latin-1"))' "$1" "$2" "$3"
}
get() { request "$1" "$2" ""; }
post() { request "$1" "" "$2"; }
# Echo a path under the root $1 exactly $2 bytes long, in NAME_MAX-sized parts.
padpath() {
local root=$1 want=$2 p=$1 last
while test $((${#p} + 203)) -lt "${want}"; do
p="${p}/$(printf '%0200d' 0)"
done
last=$((want - ${#p} - 1))
test "${last}" -ge 1 || fail "cannot build a ${want}-byte path under ${root}"
printf '%s/%0*d' "${p}" "${last}" 0
}
# Fail unless the error page renders the message opening with $1 about project
# path $2, clipped to msgmax. error.html puts ${error} alone on its own line.
check_clipped() {
local prefix=$1 full=$((${#1} + ${#2})) page line shown
test "${full}" -gt "${msgmax}" ||
fail "a ${full}-byte message fits in ${msgmax}: it cannot show a clip"
page=$(get "${port}" /server/error.html | tr -d '\r')
line=$(printf '%s\n' "${page}" | grep "^${prefix}") || {
shown=$(printf '%s\n' "${page}" | grep '^Unable to ' | cut -c1-80 || true)
fail "the error page reports '${shown:-nothing}', want: ${prefix}"
}
test "${#line}" -eq "${msgmax}" ||
fail "the message rendered ${#line} bytes, want ${full} clipped to ${msgmax}"
}
url=$(start)
port=$(portof "${url}")
srv=$(srvpid)
test -n "${srv}" || fail "htsserver did not report its pid"
# Every request body is gated by the session id (78_webhttrack-sid.test).
sid=$(get "${port}" /server/index.html |
sed -n 's/.*name="sid" value="\([0-9a-f]*\)".*/\1/p' | head -1)
test "${#sid}" -eq 32 || fail "did not scrape a 32-hex sid (got '${sid}')"
save() {
post "${port}" "sid=${sid}&command=httrack&command_do=save&winprofile=$2&path=$1&projname=$3"
}
# A project path too long to be created at all.
path=$(padpath "${base}/nodir" 1098)
save "${path}" x p >/dev/null
alive "${srv}" ||
fail "an over-long project path crashed the server: $(cat "${log}")"
check_clipped "Unable to create the directory structure in " "${path}/p"
# A creatable project path whose hts-cache is not a directory, so the init file
# under it cannot be opened.
path=$(padpath "${base}/isdir" 993)
mkdir -p "${path}/p"
ln -s /dev/null "${path}/p/hts-cache"
save "${path}" x p >/dev/null
alive "${srv}" || fail "an unwritable init file crashed the server: $(cat "${log}")"
check_clipped "Unable to create the init file " "${path}/p"
# The init file opens, then the write fails: the profile is past both stdio's
# buffer, so fwrite() reaches the descriptor, and the server's file-size limit.
path=$(padpath "${base}/full" 993)
mkdir -p "${path}/p/hts-cache"
profile=$(printf '%01024d' 0)
profile=${profile}${profile}${profile}${profile}
profile=${profile}${profile}${profile}${profile}
test "${#profile}" -eq 16384 || fail "built a ${#profile}-byte profile, want 16384"
save "${path}" "${profile}" p >/dev/null
alive "${srv}" || fail "a short init file write crashed the server: $(cat "${log}")"
written=$(wc -c <"${path}/p/hts-cache/winprofile.ini" | tr -d ' ')
test "${written}" -lt "${#profile}" ||
fail "the file-size limit did not stop the write (${written} bytes landed)"
check_clipped "Unable to write ${#profile} bytes in the the init file " "${path}/p"
get "${port}" /server/index.html | grep -q '200 OK' ||
fail "the server stopped answering after the refused saves"
echo "PASS"

View File

@@ -0,0 +1,196 @@
#!/bin/bash
#
# An unchecked box posts nothing and the stored "1" survives: hence a hidden
# companion per box, plus a disabling flag for the default-on ones.
set -euo pipefail
testdir=$(cd "$(dirname "$0")" && pwd)
distdir=${top_srcdir:-$(cd "${testdir}/.." && pwd)}
distdir=$(cd "${distdir}" && pwd)
# shellcheck source=tests/testlib.sh
. "${testdir}/testlib.sh"
fail() {
echo "FAIL: $*" >&2
exit 1
}
command -v htsserver >/dev/null || fail "no htsserver in PATH"
python=$(find_python) || {
echo "python3 not found; skipping" >&2
exit 77
}
work=$(mktemp -d "${TMPDIR:-/tmp}/webhttrack_checkbox.XXXXXX") || fail "no tmpdir"
srvlog=$(mktemp)
srv=
srvpid=
cleanup() {
# htsserver keeps SIGTERM ignored across its exec, so only -9 reaps it.
test -z "${srvpid}" || kill -9 "${srvpid}" 2>/dev/null || true
test -z "${srv}" || kill -9 "${srv}" 2>/dev/null || true
wait "${srv}" 2>/dev/null || true # absorb bash's async "Killed" notice
rm -rf "${work}" "${srvlog}"
}
trap cleanup EXIT HUP INT QUIT PIPE TERM
# An isolated HOME keeps a stray ~/.httrack.ini out of the served settings.
sport=$("${python}" -c 'import socket
s = socket.socket()
s.bind(("127.0.0.1", 0))
print(s.getsockname()[1])
s.close()')
(
trap '' TERM TTOU
export HOME="${work}"
exec htsserver "${distdir}/" --port "${sport}" >"${srvlog}" 2>&1
) &
srv=$!
for _ in $(seq 1 40); do
url=$(sed -n 's/^URL=//p' "${srvlog}") && test -n "${url}" && break
kill -0 "${srv}" 2>/dev/null || break
sleep 0.25
done
test -n "${url:-}" || fail "htsserver did not start: $(cat "${srvlog}")"
srvpid=$(sed -n 's/^PID=//p' "${srvlog}") # absent on Windows
"${python}" - "${url}" "${distdir}" <<'PY' || fail "checkbox clearing is broken (see above)"
import glob, os, re, sys, urllib.parse, urllib.request
url, srcdir = sys.argv[1].rstrip("/"), sys.argv[2]
rc = 0
def check(ok, what):
global rc
print(("ok: " if ok else "FAIL: ") + what)
if not ok:
rc = 1
# cache/cache2 are left to the separate rework of the dead "Cache" seed.
skip = {"cache", "cache2"}
page_of = {}
boxes = 0
for path in sorted(glob.glob(os.path.join(srcdir, "html", "server", "option*.html"))):
page = open(path, "rb").read().decode("latin-1")
cleared = dict((m.group(1), m.start()) for m in
re.finditer(r'<input type="hidden" name="([^"]+)" value="">', page))
for m in re.finditer(r'<input type="checkbox" name="([^"]+)"', page):
boxes += 1
if m.group(1) in skip:
continue
page_of[m.group(1)] = os.path.basename(path)
at = cleared.get(m.group(1))
check(at is not None and at < m.start(), "%s: %s is cleared before it is drawn"
% (os.path.basename(path), m.group(1)))
# Control: a regex that stopped matching would leave nothing to assert on.
check(boxes >= 25, "the option pages were scanned (%d checkboxes)" % boxes)
sid = None
def get(path):
# The UI is served ISO-8859-1, so decode, do not assume UTF-8.
return urllib.request.urlopen(url + path, timeout=20).read().decode("latin-1")
def post(fields):
global sid
if sid is None:
m = re.search(r'name="sid" value="([0-9a-f]+)"', get("/server/index.html"))
if m is None:
sys.exit("no session id in server/index.html")
sid = m.group(1)
body = "&".join("%s=%s" % (k, urllib.parse.quote(v)) for k, v in
[("sid", sid)] + fields)
req = urllib.request.Request(url + "/server/step4.html",
data=body.encode("latin-1"), method="POST")
return urllib.request.urlopen(req, timeout=20).read().decode("latin-1")
def textarea(page, name):
m = re.search(r'<textarea name="%s".*?>(.*?)</textarea>' % name, page, re.S)
if m is None:
sys.exit("no %s textarea in the rendered step4.html" % name)
return m.group(1)
def checked(page, name):
return re.search(r'name="%s"[ \t]*checked' % name, get("/server/" + page)) is not None
# What each state puts on the command line (None: nothing) and its profile key.
# Only a value-taking alias can carry a disabling flag: -I and -%I are "single"
# (htsalias.c), so a --index=0 would read back as the bare enabling flag.
BOXES = [
# field profile key set cleared
("parseall", "ParseAll", "--near", None),
("link", "Near", "--test", None),
("testall", "Test", "--extended-parsing", None),
# htmlfirst emits --priority=7, which the scan-priority list also emits.
("htmlfirst", "HTMLFirst", None, None),
("errpage", "NoErrorPages", "--generate-errors=0", None),
("external", "NoExternalPages", "--replace-external", None),
("hidepwd", "NoPwdInPages", "--disable-passwords", None),
("hidequery", "NoQueryStrings", "--include-query-string=0", None),
("nopurge", "NoPurgeOldFiles", "--purge-old=0", None),
("windebug", None, "--debug-headers", None),
("ka", "KeepAlive", "--keep-alive", None),
("remt", "RemoveTimeout", "--host-control=1", None),
("rems", "RemoveRateout", "--host-control=2", None),
("cookies", "Cookies", None, "--cookies=0"),
("parsejava", "ParseJava", None, "--parse-java=0"),
("updhack", "UpdateHack", "--updatehack", None),
("urlhack", "URLHack", "--urlhack", None),
("keepwww", "KeepWww", "--keep-www-prefix", None),
("keepslashes", "KeepSlashes", "--keep-double-slashes", None),
("keepqueryorder", "KeepQueryOrder", "--keep-query-order", None),
("toler", "TolerantRequests", "--tolerant", None),
("http10", "HTTP10", "--http-10", None),
("warc", "Warc", "--warc", None),
("changes", "Changes", "--changes", None),
("norecatch", "NoRecatch", "--do-not-recatch", None),
("logf", "Log", "--single-log", None),
("index", "Index", None, None),
("index2", "WordIndex", "--search-index", None),
("ftpprox", "UseHTTPProxyForFTP", "--httpproxy-ftp", None),
]
for field, key, when_set, when_clear in BOXES:
for value, want, unwanted, stored in (("on", when_set, when_clear, "1"),
("", when_clear, when_set, "0")):
rendered = post([(field, value)])
cmd = textarea(rendered, "command").split()
label = "%s %s" % (field, "set" if value else "cleared")
if want:
check(want in cmd, "%s emits %s" % (label, want))
if unwanted:
check(unwanted not in cmd, "%s drops %s" % (label, unwanted))
if key:
# The Windows GUI reads the same option out of the profile.
check("%s=%s" % (key, stored) in textarea(rendered, "winprofile").splitlines(),
"%s writes %s=%s" % (label, key, stored))
check(checked(page_of[field], field) == bool(value),
"%s draws %s box" % (label, "a ticked" if value else "an empty"))
missing = sorted(set(page_of) - set(b[0] for b in BOXES))
check(not missing, "every checkbox is exercised (missing %s)" % missing)
# A browser posts the companion and the ticked box under one name, so the fix
# rests on the body loop keeping the last value; both orders pin that down.
cmd = textarea(post([("cookies", ""), ("cookies", "on")]), "command").split()
check("--cookies=0" not in cmd, "cookies=&cookies=on keeps cookies set")
check(checked("option8.html", "cookies"), "cookies=&cookies=on draws a ticked box")
cmd = textarea(post([("cookies", "on"), ("cookies", "")]), "command").split()
check("--cookies=0" in cmd, "cookies=on&cookies= clears cookies")
check(not checked("option8.html", "cookies"), "cookies=on&cookies= draws an empty box")
sys.exit(rc)
PY
# A leaked htsserver wedges the parallel harness behind a green log.
cleanup
! kill -0 "${srv}" 2>/dev/null || fail "htsserver ${srv} survived"
echo "PASS"

View File

@@ -0,0 +1,181 @@
#!/bin/bash
#
# fopen() succeeds on a directory on POSIX and reading one never reaches EOF; a
# FIFO blocks in fopen() instead. Either used to wedge the single-threaded server.
set -euo pipefail
testdir=$(cd "$(dirname "$0")" && pwd)
distdir=${top_srcdir:-$(cd "${testdir}/.." && pwd)}
distdir=$(cd "${distdir}" && pwd)
fail() {
echo "FAIL: $*" >&2
exit 1
}
command -v htsserver >/dev/null || fail "no htsserver in PATH"
command -v python3 >/dev/null || {
echo "python3 not found; skipping" >&2
exit 77
}
srv=
log=$(mktemp)
base=$(mktemp -d)
cleanup() {
test -z "${srv}" || kill -9 "${srv}" 2>/dev/null || true
rm -f "${log}"
rm -rf "${base}"
}
trap cleanup EXIT HUP INT QUIT PIPE TERM
# First line only. A "| head -1" would close the pipe early and, under pipefail,
# SIGPIPE the producer into a spurious failure.
firstline() { echo "${1%%$'\n'*}"; }
freeport() {
python3 -c 'import socket
s = socket.socket()
s.bind(("127.0.0.1", 0))
print(s.getsockname()[1])
s.close()'
}
# Echo the announced URL. Runs in a command substitution, so it is a
# subshell and cannot export the pid: the caller reads it back with srvpid.
start() {
local port url
port=$(freeport)
: >"${log}"
(
trap '' TERM TTOU
exec htsserver "${distdir}/" --port "${port}" >"${log}" 2>&1
) &
for _ in $(seq 1 40); do
url=$(sed -n 's/^URL=//p' "${log}" 2>/dev/null) && test -n "${url}" && break
sleep 0.25
done
test -n "${url:-}" || fail "htsserver did not come up: $(cat "${log}")"
echo "${url}"
}
# The server reports its own pid; the aliveness assertion below hangs off it.
srvpid() { firstline "$(sed -n 's/^PID=//p' "${log}")"; }
alive() { kill -0 "$1" 2>/dev/null; }
portof() { echo "${1##*:}" | tr -d /; }
# GET the path $2 from 127.0.0.1:$1, or POST the body $3 to / when $2 is empty.
# The socket timeout is what makes an unfixed tree fail rather than wedge CI.
request() {
python3 -c 'import socket, sys
port, path, body = int(sys.argv[1]), sys.argv[2], sys.argv[3]
if path:
req = "GET %s HTTP/1.0\r\nHost: 127.0.0.1\r\n\r\n" % path
else:
req = ("POST / HTTP/1.0\r\nHost: 127.0.0.1\r\n"
"Content-type: application/x-www-form-urlencoded\r\n"
"Content-length: %d\r\n\r\n%s" % (len(body), body))
s = socket.create_connection(("127.0.0.1", port), 10)
s.settimeout(15)
s.sendall(req.encode())
out = b""
try:
while True:
b = s.recv(65536)
if not b:
break
out += b
except socket.timeout:
sys.stderr.write("timed out after %d bytes\n" % len(out))
sys.exit(9)
s.close()
sys.stdout.write(out.decode("latin-1"))' "$1" "$2" "$3"
}
# GET $2, or fail with $3 when the reply never comes.
get() {
local out
out=$(request "$1" "$2" "") || fail "$2 did not answer: $3"
echo "${out}"
}
post() { request "$1" "" "$2"; }
url=$(start)
port=$(portof "${url}")
srv=$(srvpid)
test -n "${srv}" || fail "htsserver did not report its pid"
# Needs no session id and no project: html/server is a directory of the install.
reply=$(get "${port}" /server/ "an unauthenticated request wedged the server")
case "${reply}" in
"HTTP/1.0 404 "*) ;;
*) fail "a GUI directory did not answer 404: $(firstline "${reply}")" ;;
esac
# Every request body is gated by the session id.
reply=$(get "${port}" /server/index.html "the GUI is not served")
sid=$(firstline "$(echo "${reply}" |
sed -n 's/.*name="sid" value="\([0-9a-f]*\)".*/\1/p')")
test "${#sid}" -eq 32 || fail "did not scrape a 32-hex sid (got '${sid}')"
mkdir -p "${base}/proj/sub"
echo "PROBEMARKER" >"${base}/proj/probe.txt"
# mkfifo and ln -s may not work on a Windows checkout, so both stay optional.
refuse=(/website/ /website/sub/)
if mkfifo "${base}/proj/fifo.txt" 2>/dev/null; then
refuse+=(/website/fifo.txt)
fi
ln -s probe.txt "${base}/proj/link.txt" 2>/dev/null || true
# step4's "save settings" flow registers the project without crawling, and that
# is what arms the /website/ root.
body="sid=${sid}&command=httrack&command_do=save&winprofile=x"
body="${body}&path=${base}&projname=proj"
post "${port}" "${body}" >/dev/null || fail "the save request did not answer"
test -f "${base}/proj/hts-cache/winprofile.ini" ||
fail "the project was not registered: $(cat "${log}")"
# Positive control: every assertion below would pass on a server serving nothing.
served() {
case "$(get "$1" "$2" "$3")" in
*PROBEMARKER*) ;;
*) fail "$3" ;;
esac
}
served "${port}" /website/probe.txt "the registered project's mirror is not served"
if test -L "${base}/proj/link.txt"; then
# The guard stats rather than lstats, so a symlinked mirror file still serves.
served "${port}" /website/link.txt "a symlink to a mirror file is not served"
fi
# /website/ is where the GUI's own "browse mirrored site" link points.
for path in "${refuse[@]}"; do
reply=$(get "${port}" "${path}" "the server wedged on ${path}")
case "${reply}" in
"HTTP/1.0 404 "*) ;;
*) fail "${path} did not answer 404: $(firstline "${reply}")" ;;
esac
done
# The accept loop is single-threaded: one spin denies every later request.
alive "${srv}" || fail "the server died: $(cat "${log}")"
served "${port}" /website/probe.txt "the server stopped answering after a directory request"
# This fopen() is reached before any guard, so its read loop must stop on ferror().
mkdir -p "${base}/proj2/hts-cache/winprofile.ini"
reply=$(post "${port}" "sid=${sid}&path=${base}&loadprojname=proj2") ||
fail "a directory winprofile.ini wedged the server"
case "${reply}" in
"HTTP/1.0 3"*) ;;
*) fail "loading a project did not redirect: $(firstline "${reply}")" ;;
esac
alive "${srv}" || fail "the server died: $(cat "${log}")"
served "${port}" /website/probe.txt "the server stopped answering after a project load"
echo "PASS"

View File

@@ -0,0 +1,37 @@
#!/bin/bash
# The two halves of an .ndx entry's URL share one buffer, so the second must be
# bounded by what the first left. The sanitizer CI build turns a regression
# here into a hard stack-buffer-overflow.
set -euo pipefail
testdir=$(cd "$(dirname "$0")" && pwd)
# shellcheck source=tests/testlib.sh
. "${testdir}/testlib.sh"
python=$(find_python) || {
echo "python3 not found; skipping" >&2
exit 77
}
dir=$(mktemp -d)
trap 'rm -rf "$dir"' EXIT
# each field fills binput's own bound, so together they run one past line[]
"$python" - "$dir/foo.ndx" <<'EOF'
import sys
with open(sys.argv[1], "w") as f:
f.write("CACHE-1.4\n")
f.write("Wed, 01 Jan 2025 00:00:00 GMT\n")
f.write("h" * 1024 + "\n" + "f" * 1024 + "\n" + "0\n")
# a trailing entry keeps the walk inside the buffer, so this exercises the
# field bound alone
f.write("tail/entry\nx\n1\n")
EOF
: >"$dir/foo.dat" # the reader only proceeds when the sibling .dat opens
proxytrack --convert "$dir/out.arc" "$dir/foo.ndx" >/dev/null 2>&1 || {
echo "FAIL: proxytrack crashed on an .ndx with two full-length URL fields" >&2
exit 1
}

246
tests/93_local-changes.test Normal file
View File

@@ -0,0 +1,246 @@
#!/bin/bash
# --changes (#714): the report must match the mirror's real deltas. Every route
# answers 200 with no validators on both passes, so the transfer signal alone
# would call the whole site changed.
set -euo pipefail
: "${top_srcdir:=..}"
testdir=$(cd "$(dirname "$0")" && pwd)
# shellcheck source=tests/testlib.sh
. "${testdir}/testlib.sh"
server=$(nativepath "${testdir}/local-server.py")
root=$(nativepath "${testdir}/server-root")
python=$(find_python) || ! echo "python3 not found; skipping" >&2 || exit 77
tmpdir=$(mktemp -d "${TMPDIR:-/tmp}/httrack_changes.XXXXXX") || exit 1
serverpid=
cleanup() {
stop_server "$serverpid"
rm -rf "$tmpdir"
}
trap cleanup EXIT HUP INT QUIT PIPE TERM
serverlog="${tmpdir}/server.log"
: >"$serverlog"
"$python" "$server" --root "$root" >"$serverlog" 2>&1 &
serverpid=$!
port=
for _ in $(seq 1 50); do
line=$(head -n1 "$serverlog" 2>/dev/null)
if test "${line%% *}" == "PORT"; then
port="${line#PORT }"
break
fi
kill -0 "$serverpid" 2>/dev/null || {
echo "server exited early: $(cat "$serverlog")"
exit 1
}
sleep 0.1
done
test -n "$port" || {
echo "could not discover server port"
exit 1
}
base="http://127.0.0.1:${port}"
which httrack >/dev/null || {
echo "could not find httrack"
exit 1
}
out="${tmpdir}/crawl"
mkdir "$out"
report="${out}/hts-changes.json"
# Purging off for the first two passes: the deleted set must be computed with
# nothing acting on it.
common=(--quiet --disable-security-limits --robots=0 --timeout=30 --changes --retries=2)
# Report field extractor: a scalar, and the files listed under a bucket.
field() { sed -n 's/.*"'"$1"'": \([a-z0-9]*\).*/\1/p' "${2:-$report}" | head -1; }
# tr: python prints CRLF on Windows, and the expectations below are LF.
listed() {
"$python" - "${2:-$report}" "$1" <<'EOF' | tr -d '\r'
import json
import sys
with open(sys.argv[1], encoding="utf-8") as fp:
report = json.load(fp)
for entry in report[sys.argv[2]]:
print(entry["file"])
EOF
}
expect() {
local label="$1" want="$2" got="$3" rep="${4:-$report}"
printf '[%s] ..\t' "$label"
test "$want" == "$got" || {
echo "FAIL: expected '${want}', got '${got}'"
cat "$rep"
exit 1
}
echo "OK"
}
# Every count must equal the length of the list it counts.
expect_counts_match() {
local label="$1" rep="${2:-$report}" bucket n
printf '[%s] ..\t' "$label"
for bucket in new changed unchanged gone; do
n=$(listed "$bucket" "$rep" | grep -c . || true)
test "$(field "$bucket" "$rep")" == "$n" || {
echo "FAIL: ${bucket} count $(field "$bucket" "$rep") but $n listed"
cat "$rep"
exit 1
}
done
echo "OK"
}
lines() { printf '%s\n' "$@" | sort; }
# A connection killed before the status line surfaces differently per platform
# (macOS truncates the file, #748; Linux leaves it), so reset.bin is asserted on
# its own and kept out of the exact lists.
listed_but_reset() { listed "$1" | grep -v '/reset\.bin$' | sort; }
# --- pass 1: nothing to compare against --------------------------------------
httrack "${common[@]}" -O "$out" --purge-old=0 "${base}/changes/index.html" \
>"${tmpdir}/log1" 2>&1
test -s "$report" || {
echo "FAIL: pass 1 wrote no ${report}"
exit 1
}
expect "pass 1 declares a first crawl" "true" "$(field first_crawl)"
# Every mirrored file: index, stable.html, moved.html, stable.bin, moved.bin,
# doomed.html, redirtarget.html, flaky.bin, coded.bin, codedstable.bin,
# reset.bin, sized.html and the redirect target it points at.
expect "pass 1 counts 13 new" "13" "$(field new)"
expect "pass 1 counts 0 changed" "0" "$(field changed)"
expect "pass 1 counts 0 gone" "0" "$(field gone)"
expect_counts_match "pass 1 counts match its lists"
grep -aq "first crawl" "${out}/hts-log.txt" || {
echo "FAIL: pass 1 log carries no first-crawl summary"
exit 1
}
host="127.0.0.1_${port}"
# A leftover from some earlier failed attempt, at a name pass 2 fetches for the
# first time: on disk, but never part of the previous mirror, so it is new.
printf 'leftover junk' >"${out}/${host}/changes/fresh.html"
# --- pass 2: two bodies move, pages appear and disappear ---------------------
httrack "${common[@]}" -O "$out" --purge-old=0 --update \
"${base}/changes/index.html" >"${tmpdir}/log2" 2>&1
expect "pass 2 is not a first crawl" "false" "$(field first_crawl)"
# The redirect behind sized.html now points at a much longer target name.
long="${host}/changes/$(printf 's%.0s' $(seq 40)).html"
expect "new is the two fresh pages and the renamed redirect target" \
"$(lines "${host}/changes/fresh.html" "${host}/changes/transient.html" "$long")" \
"$(listed new | sort)"
expect "gone is doomed.html and the old redirect target" \
"$(lines "${host}/changes/doomed.html" "${host}/changes/s.html")" \
"$(listed gone | sort)"
# index.html changes because its link list does; moved.* change their bodies.
# coded.bin arrives gzipped and direct-to-disk, so its previous copy has to be
# sampled before the decoded temp is renamed over it; it is the same length on
# both passes, so only a digest separates it from codedstable.bin.
expect "changed is exactly the four moved resources" \
"$(lines "${host}/changes/coded.bin" "${host}/changes/index.html" \
"${host}/changes/moved.bin" "${host}/changes/moved.html")" \
"$(listed_but_reset changed)"
# Re-served byte for byte, 200, no Last-Modified, no ETag. flaky.bin gets there
# through a failed transfer and a retry, so its file is notified twice;
# redirtarget.html arrives behind a 302. sized.html has a byte-identical
# payload, but the rewritten link to the renamed redirect target makes the file
# on disk change length.
expect "unchanged is exactly the six stable resources" \
"$(lines "${host}/changes/codedstable.bin" "${host}/changes/flaky.bin" \
"${host}/changes/redirtarget.html" "${host}/changes/sized.html" \
"${host}/changes/stable.bin" "${host}/changes/stable.html")" \
"$(listed_but_reset unchanged)"
# reset.bin never completes a transfer, so it drops out of new.lst with its
# mirrored copy still there. Whatever else it is, it is not a deletion.
expect "a failed re-fetch is never reported gone" "" \
"$(listed gone | grep '/reset\.bin$' || true)"
expect_counts_match "pass 2 counts match its lists"
# Every mirrored file appears once and only once, across all four buckets.
printf '[each file in exactly one bucket] ..\t'
dupes=$( (
listed new
listed changed
listed unchanged
listed gone
) | sort | uniq -d)
test -z "$dupes" || {
echo "FAIL: listed in more than one bucket: ${dupes}"
exit 1
}
echo "OK"
# The log summary is the report, rendered: it must not drift from the counts.
summary="Change report: $(field new) new, $(field changed) changed, $(field unchanged) unchanged, $(field gone) gone (hts-changes.json)"
printf '[the log summary matches the counts] ..\t'
grep -aqF "$summary" "${out}/hts-log.txt" || {
echo "FAIL: no '${summary}' in the log"
grep -a "Change report" "${out}/hts-log.txt" || true
exit 1
}
echo "OK"
expect "the report says nothing was purged" "false" "$(field purged)"
printf '[reporting gone did not delete it] ..\t'
test -f "${out}/${host}/changes/doomed.html" || {
echo "FAIL: doomed.html was deleted with --purge-old=0"
exit 1
}
echo "OK"
printf '[a failed transfer keeps its file] ..\t'
test -f "${out}/${host}/changes/reset.bin" || {
echo "FAIL: reset.bin lost its previous copy"
exit 1
}
echo "OK"
# --- pass 3: the same run with purging on ------------------------------------
httrack "${common[@]}" -O "$out" --update "${base}/changes/index.html" \
>"${tmpdir}/log3" 2>&1
expect "the report says files were purged" "true" "$(field purged)"
expect "gone is the page that only pass 2 linked" \
"${host}/changes/transient.html" "$(listed_but_reset gone)"
printf '[a purged file is off disk] ..\t'
test ! -f "${out}/${host}/changes/transient.html" || {
echo "FAIL: transient.html survived --purge-old"
exit 1
}
echo "OK"
# --- cache off: the report degrades, it does not lie -------------------------
nocache="${tmpdir}/nocache"
mkdir "$nocache"
ncreport="${nocache}/hts-changes.json"
for pass in 1 2; do
httrack "${common[@]}" -O "$nocache" -C0 "${base}/changes2/index.html" \
>"${tmpdir}/lognc${pass}" 2>&1
done
expect "with no cache first_crawl is undetermined, not guessed" "null" \
"$(field first_crawl "$ncreport")" "$ncreport"
# The bodies on disk are still compared, so a constant one is not called
# changed. index.html is: a parsed page is compared against the cached payload,
# and there is no cache, which is the degradation the report declares.
expect "the constant body is still recognised as unchanged" \
"${host}/changes2/stable2.bin" "$(listed unchanged "$ncreport")" "$ncreport"
expect "the moving body is still recognised as changed" \
"$(lines "${host}/changes2/index.html" "${host}/changes2/ticker2.bin")" \
"$(listed changed "$ncreport" | sort)" "$ncreport"
expect_counts_match "cache-off counts match its lists" "$ncreport"
printf '[every entry carries its size] ..\t'
"$python" - "$ncreport" <<'EOF' || exit 1
import json
import sys
with open(sys.argv[1], encoding="utf-8") as fp:
report = json.load(fp)
for bucket in ("new", "changed", "unchanged"):
for entry in report[bucket]:
if "size" not in entry:
sys.exit("FAIL: %s has no size in %s" % (entry["file"], bucket))
EOF
echo "OK"

View File

@@ -30,6 +30,7 @@ TEST_EXTENSIONS = .test
TEST_LOG_COMPILER = $(BASH)
TESTS = \
00_runnable.test \
01_engine-changes.test \
01_engine-charset.test \
01_engine-cmdline.test \
01_engine-cmdline-split.test \
@@ -183,6 +184,13 @@ TESTS = \
83_webhttrack-argescape.test \
84_webhttrack-mirror-verbatim.test \
85_webhttrack-projpath.test \
86_local-proxytrack-cache-longfields.test
86_local-proxytrack-cache-longfields.test \
87_local-proxytrack-nodate.test \
88_local-proxytrack-badmtime.test \
89_webhttrack-error-overflow.test \
90_webhttrack-checkbox-clear.test \
91_webhttrack-directory.test \
92_local-proxytrack-ndx-fields.test \
93_local-changes.test
CLEANFILES = check-network_sh.cache

View File

@@ -1571,6 +1571,167 @@ class Handler(SimpleHTTPRequestHandler):
if self.command != "HEAD":
self.wfile.write(body)
# --changes (#714). Every route answers 200 with no validators, so the
# transfer signal alone would call the whole site changed on pass 2; only a
# payload comparison can tell stable.html and stable.bin from moved.*.
def route_changes_index(self):
links = "".join(
'\t<a href="%s">%s</a>\n' % (name, name)
for name in (
"stable.html",
"moved.html",
"stable.bin",
"moved.bin",
"redir.html",
"flaky.bin",
"coded.bin",
"codedstable.bin",
"sized.html",
"reset.bin",
)
)
seen = self.refetch_pass()
if seen == 1:
links += '\t<a href="doomed.html">doomed</a>\n'
else:
links += '\t<a href="fresh.html">fresh</a>\n'
# transient.html appears on pass 2 only, so pass 3 has a deletion of
# its own to purge.
if seen == 2:
links += '\t<a href="transient.html">transient</a>\n'
self.send_html(links)
def route_changes_stable(self):
self.refetch_pass()
self.send_raw(b"<html><body><p>CHANGES-STABLE</p></body></html>", "text/html")
def route_changes_moved(self):
pass1 = self.refetch_pass() == 1
self.send_raw(
b"<html><body><p>CHANGES-MOVED-V%d</p></body></html>" % (1 if pass1 else 2),
"text/html",
)
def route_changes_stable_bin(self):
self.refetch_pass()
self.send_raw(
b"CHANGES-STABLE-BIN\n" + b"\x00\x01\x02\xff" * 512,
"application/octet-stream",
)
def route_changes_moved_bin(self):
pass1 = self.refetch_pass() == 1
self.send_raw(
b"CHANGES-MOVED-BIN-V%d\n" % (1 if pass1 else 2)
+ b"\x03\x02\x01\xfe" * 512,
"application/octet-stream",
)
def route_changes_doomed(self):
self.refetch_pass()
self.send_raw(b"<html><body><p>CHANGES-DOOMED</p></body></html>", "text/html")
def route_changes_fresh(self):
self.refetch_pass()
self.send_raw(b"<html><body><p>CHANGES-FRESH</p></body></html>", "text/html")
def route_changes_redir(self):
self.send_response(302)
self.send_header("Location", "/changes/redirtarget.html")
self.send_header("Content-Length", "0")
self.end_headers()
def route_changes_redirtarget(self):
self.refetch_pass()
self.send_raw(b"<html><body><p>CHANGES-REDIR</p></body></html>", "text/html")
# Direct-to-disk, and the first body fetch of each pass dies mid-transfer:
# the retry notifies the same file a second time, so the accumulator has to
# keep the pre-run sample from the first notify.
FLAKY = b"CHANGES-FLAKY\n" + b"\x07\x06\x05\x04" * 512
# A page whose payload never changes, linking through a redirect whose
# target is renamed between passes: the rewritten link makes the file on
# disk change length while the payload is byte-identical. Only a payload
# comparison can call it unchanged.
SIZED_SHORT = "s.html"
SIZED_LONG = "s" * 40 + ".html"
def route_changes_sized(self):
self.refetch_pass()
self.send_raw(
b'<html><body><a href="sizedredir.html">t</a></body></html>',
"text/html",
)
def route_changes_sizedredir(self):
target = self.SIZED_SHORT if self.refetch_pass() == 1 else self.SIZED_LONG
self.send_response(302)
self.send_header("Location", "/changes/" + target)
self.send_header("Content-Length", "0")
self.end_headers()
def route_changes_sizedtarget(self):
self.refetch_pass()
self.send_raw(b"<html><body><p>SIZED-TARGET</p></body></html>", "text/html")
# Content-Encoding on a direct-to-disk body: the decoded temp is renamed
# over the mirror file, so the previous copy has to be sampled before the
# rename. Same length on both passes, so only a digest can tell them apart.
def route_changes_coded(self):
pass1 = self.refetch_pass() == 1
body = b"CHANGES-CODED-V%d\n" % (1 if pass1 else 2) + b"\x11\x22" * 1024
self.send_coded(self.gzipped(body), "application/octet-stream")
# Control: same coding, same bytes on both passes.
def route_changes_coded_stable(self):
self.refetch_pass()
body = b"CHANGES-CODED-STABLE\n" + b"\x33\x44" * 1024
self.send_coded(self.gzipped(body), "application/octet-stream")
def route_changes_flaky(self):
if self.refetch_pass() % 2 == 1:
self.send_truncated(self.FLAKY, "application/octet-stream")
else:
self.send_raw(self.FLAKY, "application/octet-stream")
def route_changes_transient(self):
self.refetch_pass()
self.send_raw(
b"<html><body><p>CHANGES-TRANSIENT</p></body></html>", "text/html"
)
# Answers once, then hangs up before sending anything on every later
# attempt, retries included: nothing is written, so the file never reaches
# new.lst while its mirrored copy stays intact.
RESET = b"CHANGES-RESET\n" + b"\x21\x22\x23\x24" * 512
def route_changes_reset(self):
if self.refetch_pass() == 1:
self.send_raw(self.RESET, "application/octet-stream")
else:
self.close_connection = True
self.connection.close()
# A second, independent mirror crawled with the cache off: nothing but the
# bytes on disk can tell stable2 from ticker2, whose body differs every
# fetch at a constant length.
def route_changes2_index(self):
self.refetch_pass()
self.send_html('\t<a href="stable2.bin">s</a>\n\t<a href="ticker2.bin">t</a>\n')
def route_changes2_stable(self):
self.refetch_pass()
self.send_raw(
b"CHANGES2-STABLE-00\n" + b"\x55\xaa" * 512, "application/octet-stream"
)
def route_changes2_ticker(self):
self.send_raw(
b"CHANGES2-TICKER-%02d\n" % (self.refetch_pass() % 100) + b"\x55\xaa" * 512,
"application/octet-stream",
)
ROUTES = {
"/cookies/entrance.php": route_entrance,
"/cookies/second.php": route_second,
@@ -1588,6 +1749,27 @@ class Handler(SimpleHTTPRequestHandler):
"/codec/bad.html": route_codec_bad,
"/codec/bin.dat": route_codec_bin,
"/codec/ae.html": route_codec_ae,
"/changes/index.html": route_changes_index,
"/changes/stable.html": route_changes_stable,
"/changes/moved.html": route_changes_moved,
"/changes/stable.bin": route_changes_stable_bin,
"/changes/moved.bin": route_changes_moved_bin,
"/changes/doomed.html": route_changes_doomed,
"/changes/fresh.html": route_changes_fresh,
"/changes/redir.html": route_changes_redir,
"/changes/redirtarget.html": route_changes_redirtarget,
"/changes/flaky.bin": route_changes_flaky,
"/changes/coded.bin": route_changes_coded,
"/changes/sized.html": route_changes_sized,
"/changes/sizedredir.html": route_changes_sizedredir,
"/changes/" + SIZED_SHORT: route_changes_sizedtarget,
"/changes/" + SIZED_LONG: route_changes_sizedtarget,
"/changes/codedstable.bin": route_changes_coded_stable,
"/changes/transient.html": route_changes_transient,
"/changes/reset.bin": route_changes_reset,
"/changes2/index.html": route_changes2_index,
"/changes2/stable2.bin": route_changes2_stable,
"/changes2/ticker2.bin": route_changes2_ticker,
"/upcodec/index.html": route_upcodec_index,
"/upcodec/mem.html": route_upcodec_mem,
"/upcodec/disk.bin": route_upcodec_disk,

View File

@@ -46,12 +46,14 @@ cat >"$stubdir/x-www-browser" <<EOF
echo "stub browser invoked with: \$1" >&2
# Also fetch an option page and require a rendered title='' tooltip: proves the
# option template expands and the \${html:} filter escapes into the attribute.
# option9 additionally proves the WARC control renders with its expanded label.
# option9 additionally proves the WARC and change-report controls render with
# their expanded labels.
opturl="\${1%/}/server/option2.html"
warcurl="\${1%/}/server/option9.html"
if body="\$(curl -fsSL --max-time 20 "\$1")" && printf '%s' "\$body" | grep -qai httrack && printf '%s' "\$body" | grep -qaF step2.html &&
opt="\$(curl -fsSL --max-time 20 "\$opturl")" && printf '%s' "\$opt" | grep -qaF "title='" &&
warc="\$(curl -fsSL --max-time 20 "\$warcurl")" && printf '%s' "\$warc" | grep -qaF 'name="warcfile"' && printf '%s' "\$warc" | grep -qaF WARC; then
warc="\$(curl -fsSL --max-time 20 "\$warcurl")" && printf '%s' "\$warc" | grep -qaF 'name="warcfile"' && printf '%s' "\$warc" | grep -qaF WARC &&
printf '%s' "\$warc" | grep -qaF 'name="changes"' && printf '%s' "\$warc" | grep -qaF hts-changes.json; then
echo PASS >"$marker"
else
echo "FAIL: unexpected response from \$1" >"$marker"