Commit Graph

9 Commits

Author SHA1 Message Date
Xavier Roche
9e29c1e159 Sitemap files are never read, so URLs nothing links to are never found (#718)
* Read sitemap files so URLs nothing links to are found

HTTrack finds URLs only by parsing links, so anything a site publishes solely
in its sitemap stayed invisible: robots.txt was already parsed, but its
Sitemap: lines were ignored and nothing else in the tree touched sitemaps.

Adds opt-in --sitemap (-%m), which probes the start host's robots.txt and
falls back to /sitemap.xml, and --sitemap-url (-%mu) for an explicit document.
Handles <urlset> and nested <sitemapindex>, plain or gzipped. Discovered URLs
enter with the full depth budget but still go through the wizard, so filters
and scope rules decide; a sitemap is not a filter bypass.

The parser reads attacker-controlled XML off the network, so it is capped on
URL count, index nesting, decompressed size and decompression ratio, and child
sitemaps must stay on the host that named them.

Closes #712

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* sitemap: distinguish the sitemapindex log line from a urlset one

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* sitemap: fix an out-of-bounds read, tighten the parser and the tests

lienrelatif() walked back from the last character of its current-path
argument without checking the path was non-empty, reading one byte before
the stack buffer. htsAddLink is the first caller to pass an empty savename,
which sitemap documents have because they are ingested rather than mirrored,
so ASan caught it on the new crawl test.

The parser drops a value whose numeric character reference decodes outside
printable ASCII, rather than leaving the reference verbatim and seeding a URL
the site never published, and classifies a document by its real root element,
so a comment naming the other one no longer flips urlset and sitemapindex.
The robots.txt line reader is bounded by the body size instead of relying on
a NUL terminator.

The self-test moves to 01_zlib-sitemap.test: MSan runs 01_engine-* only,
because an uninstrumented libz floods it with false positives.

Tests gain the assertions the earlier ones were missing: which of the
robots.txt route and the /sitemap.xml fallback was taken, that the sitemap
documents stay out of the mirror, that the off-host child sitemap is refused,
the sitemapindex nesting cap, the per-document URL cap at its production
value, and copy_htsopt coverage for the two new fields.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* sitemap: state what the decompression cap actually binds on

deflate tops out near 1032:1, so hts_codec_maxout never binds before the
64 MiB cap; the old comment implied a ratio guard that cannot fire.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* sitemap: a child sitemap is a fetch, so filters and robots.txt must gate it

An adversarial review found that a <sitemapindex> <loc>, and a robots.txt
Sitemap: line, went straight to hts_record_link: the request went out even
when a -* rule or robots.txt Disallow covered it. Only the <urlset> half ran
through the wizard. Gate the document itself on the filters and on
robots.txt, which is all that can apply: the wizard proper wants a referring
link, and its up/down travel rules would judge a child sitemap against the
parent sitemap's own directory. The robots.txt probe is exempt, being the
request that fetches the rules.

A 301 also used to end ingestion silently, since the engine re-queues the
target as a fresh link that carried no sitemap marking. That hit any site
redirecting http to https. The marking now follows the redirect.

The "N URL(s) added" counter reported what the scanner emitted rather than
what was taken, which hid both of the above; it now reads "N of M". The
fallback to /sitemap.xml keys on the same corrected count, so a robots.txt
whose only Sitemap: line is off-host or filtered still falls back. Root
classification skips a UTF-8 BOM and an XML namespace prefix, and the doc
list is cleared when a mirror starts rather than only when it ends.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* sitemap: gate each fetch by who asked for it, and anchor travel on the start URL

A nine-agent review found a scope escape: the sitemap document was its own
`premier`, so the wizard measured travel from wherever the site chose to put
its sitemap. A root /sitemap.xml therefore widened a /deep/dir/ crawl to the
whole host. The ingester now points the wizard at the crawl's own start link
and lets each seeded URL become its own anchor, which is what a command-line
seed gets.

Robots handling was both mistimed and undifferentiated. The Sitemap: lines are
now collected by robots_parse, on the same body in the same fetch, and acted on
after the parsed rules are installed rather than before; and the decision comes
from a new hts_robots_forbids extracted out of the wizard, so the sitemap path
inherits the -s1 filters-win override instead of a stricter hand-rolled check.
The four fetches are no longer treated alike: a sitemap the user names is user
intent, one the site declares invites the fetch, only the guessed /sitemap.xml
obeys a Disallow, and the URLs listed inside stay fully gated.

Also: hts_unescapeEntities replaces the private entity decoder, whose guard
tests and fuzz corpus it silently forfeited; hts_codec_head replaces hts_zhead,
which is only defined under HTS_USEZLIB; the composed URL buffer now fits two
maximal components plus a scheme, which a 2046-byte --sitemap-url reached; the
bounded search is promoted to htstools as hts_memstr; and the live state moves
from httrackp into htsoptstate, leaving two installed fields rather than three.

Tests gain the scope escape, the three robots cases, a cap-boundary control,
the handler invocation count and a compression-bomb decode. Every one was
checked against a deliberately broken build.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* sitemap: add the fuzz harness the parser was missing, and drop truncated Sitemap: lines

The file header called the scanner fuzzable while fuzz/ registered ten
harnesses and none for it. fuzz-sitemap feeds it raw XML, gzip-framed bodies
and truncated streams off a heap copy of exactly the input size, so an overread
is an ASan report rather than a quiet pass, with a four-file seed corpus.
60000 runs clean under ASan+UBSan.

robots_parse now drops a Sitemap: line that filled its scratch buffer instead
of handing on the half URL it was truncated to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* fuzz: keep only the four sitemap seed inputs

A libFuzzer run writes its finds into the first corpus directory, and 191 of
them were committed with the harness.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* selftest: bound the sitemap document builders' snprintf accumulation

snprintf returns the length it wanted to write, so accumulating it blind
lets the next offset and size argument walk past the buffer. Guard each
step the way the argv builder above already does, and give the per-URL
loop a real remaining-space bound instead of a fixed 33.

Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* sitemap: date the new files 2026

The headers were copied from an existing file and kept its 1998 year.

Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* sitemap: keep the ingestion state out of htsoptstate

htsoptstate is embedded by value as httrackp.state, so a field at its tail
shifts every httrackp member declared after it: an offsetof probe put
warc_file at 141752 on master and 141760 on the branch. Move the pointer to
httrackp's own tail, where every existing offset holds and copy_htsopt still
ignores it.

Also renumber the crawl test to 89, master having taken 87 and 90, and give
the new option8 checkbox the hidden companion that 90_webhttrack-checkbox-clear
requires, plus its row in that test's table.

Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: renumber the sitemap crawl test to 95

Master took 88 through 93 and #720 claims 94.

Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* sitemap: realign the htsopt.h comments after the single-file merge

Master's longer LLint declarator moved the block's comment column.

Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* sitemap: translate the new GUI strings into the remaining 28 locales (#738)

The feature PR added the four LANG_SITEMAP* entries to lang.def with English
and Francais only; every other locale fell back to English in the WebHTTrack
form. Each file is written in its own declared charset.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 07:59:59 +02:00
Xavier Roche
0774d47d2f No single-file HTML output with inlined assets (#720)
* Single-file output is MHT, which browsers no longer open

Add --single-file (-%Z): after the mirror completes, rewrite every saved
page in place with its stylesheets, scripts, images and fonts embedded as
data: URIs, while links between pages stay relative. The mirror remains a
browsable tree and each page also stands alone.

This cannot reuse the -%M path: MHT streams a MIME part per file as it is
saved, but a data: URI needs the asset's bytes when the page is written,
and pages are normally saved before their assets are fetched. The new pass
runs over the finished tree at the tail of httpmirror(), after the update
purge.

Audio, video, page-to-page links and anything over --single-file-max-size
(10 MB default) keep an ordinary link. References carrying a scheme are
skipped, which covers data: and makes a second --update run a no-op.
Resolution is clamped to the mirror root: the HTML is hostile input.

htsopt.h gains two tail-appended fields; VERSION_INFO is untouched.

Closes #713

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>

* tests: renumber the single-file test to 87, run the self-test last

Master took 82 through 85 while this branch was open. The engine self-test
also moves to the end of the script: it and the crawl assertions cover
different ground, and failing first hid the crawl half.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>

* Fix the parser and path bugs review found, and cover them

sf_relative_from() indexed one byte past from_dir's terminator whenever the
page's own directory was a prefix of the asset path, which is the ordinary
mirror layout: an out-of-bounds read that also miscounted the ../ prefix and
emitted a broken link. Two review agents hit the same ASan trace; the branch
had no test because every nested asset in the fixture was under the cap.

Also from review: an escaped quote no longer ends a CSS string early (which
exposed its contents to the url()/@import scanner), @import url(...) now
inlines like the quoted form, a raw-text element ends only on a real end tag
rather than any prefix of one, an over-wide tag is copied through by the same
quote-aware scan instead of a second quote-blind one, <!--> is an empty
comment, whitespace before a tag's > survives, and a page cannot inline more
than SINGLEFILE_MAX_PAGE_SIZE, which bounds the multiplicative @import
fan-out. A failed encode no longer leaves a payload-less data: prefix behind,
and base_dir is matched against the root on a component boundary.

sf_readfile now wraps a new readfile2_utf8() rather than being a fifth copy
of the readfile family. Docs place --single-file beside -%M instead of
implying it supersedes it: MHT stays the better container, this wins on
opening anywhere.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>

* singlefile: tighten two comments

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>

* tests: cover the option plumbing review found untested

copy_htsopt's two new fields, including the >0 guard that must not let an
unset source clear the target's default; -%Z0; and a rejected
--single-file-max-size argument falling back to the built-in cap.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>

* tests: stand the single-file GUI check down where htsserver is absent

The Windows job builds httrack.exe only, and reaches this file through its
*_local-*.test glob, so requiring htsserver failed the whole test there even
though every crawl assertion had passed. Skip that half instead, and run the
engine self-test before it so the skip cannot swallow it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>

* tests: check the sprintfbuff master just made warn_unused_result

#722 gave slprintfbuff the attribute, so the over-wide-tag fixture's call
became the one warning this branch adds over master's baseline. Handle it the
way that PR's own self-test code does; the buffer cannot truncate here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>

* single-file: translate the new GUI strings into the remaining 28 locales (#739)

The feature PR added the four LANG_SINGLEFILE* entries to lang.def with English
and Francais only; every other locale fell back to English in the WebHTTrack
form. Each file is written in its own declared charset.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* single-file: restore the mirror file mode, and give the guards real coverage

The rewrite spools to a .sfnew temp and renames it over the page, which
bypasses filecreate() and with it the HTS_ACCESS_FILE chmod the engine puts on
every other mirrored file. Under a restrictive umask the pages came out 0600
while their assets stayed 0644, so a mirror served by a webserver or shared
with a group lost read access on exactly the pages. chmod the spool before the
rename.

Two guards had no coverage: deleting the scheme/data: check in sf_resolve left
both the self-test and the crawl test passing, because the fixtures resolved to
paths that were absent either way, and the per-page inline budget was never
exercised. The fixtures now plant a file where each guard's removal would land
the walk, and a self-importing stylesheet measures the budget against a
large-budget control. Charging that budget after the nested rewrite instead of
before let an @import chain spend what its ancestors had already claimed and
drove it negative; charge it up front and refund on failure.

The attribute table missed the lazy-loading attributes hts_detect[] already
downloads, so a modern page inlined almost nothing: add data-src, data-srcset,
lowsrc, object@data and embed@src, and record why the rest stay links.

The new web GUI checkbox had no hidden companion input, so it could be ticked
but never cleared, which is what #725 fixed for every other box. Master's test
90 catches it once the branch merges.

Adds a libFuzzer harness over the rewriter, since it re-serializes hostile
HTML, and renames the crawl test to 91 now that master holds 87 and 90.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* single-file: document the inlined-stylesheet limitation in the CLI guide

Recorded in htssinglefile.h already; the user-facing guide is where someone
raising --single-file-max-size will look.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: move the Windows note down to the gate it explains

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* ci: register the single-file test's Windows skip

Its GUI half needs htsserver, which the Windows job does not build, so the test
now exits 77 there instead of reporting a pass for assertions it never ran. The
skip list is pinned, so it has to be declared. Both Windows jobs reported
fail=0; only the list check was red.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <xroche@gmail.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 07:17:13 +02:00
Xavier Roche
7e2209d4cc No way to see what changed between two mirrors (#721)
* Report what a crawl changed against the previous mirror (--changes)

--update already knows which resources were new, which changed and which the
server called unchanged, and throws it away: the flags reach file_notify() and
go no further than a log line, while deletions exist only as a side effect of
purging. --changes (-%d) keeps all of it and writes hts-changes.json plus a
one-line summary in the log.

"Changed" means the bytes differ, not that the server re-sent the resource.
Comparing the mirrored files directly would not work: HTTrack stamps every
parsed page with the crawl date via the footer, so those bytes differ on every
run. Payloads are compared instead, the previous one coming from the cache for
parsed pages and from the local copy sampled just before it is overwritten for
everything else.

The mirror-relative path, not the URL, is the accumulator's key, so a redirect
and its target that share a save name are one entry; and only the first notify
for a file samples its pre-run state, so a retried transfer is not counted
twice. What counts as already mirrored comes from the previous run's file
index rather than from the file's presence on disk: a partial left by this
crawl's own failed attempt is on disk but was never part of the previous
mirror.

The deleted set is now computed whether or not purging is enabled; unlinking
still happens only under --purge-old.

Closes #714

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>

* changes: never leave a stale report behind

A crawl that mirrored nothing created no accumulator, so hts_changes_close_opt
returned without writing and the previous run's report stayed on disk as if it
described this one. Write it whenever --changes is on. The no-data rollback is
the deliberate exception, and is now documented: it restores the previous cache
generation, so leaving the matching report alone is the consistent behaviour.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>

* changes: drop em dashes from the format page

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>

* changes: fix the review's three findings

The size shortcut compared rendered on-disk sizes even for parsed pages, whose
payload digests describe something else entirely, so it decided the outcome
before the payload comparison could run: a page whose payload never changed but
whose rewritten links moved read as changed. It now only applies when both
digests describe the file on disk.

file_notify() reaches the accumulator from the FTP download thread as well as
the main one, and the lazy allocation, the coucal write and the entries realloc
were all unguarded. Every entry point now takes a mutex, and the HTML hook does
its cache read before taking it, since that read can itself re-enter
file_notify() and move the array.

Two fixtures cover what nothing did: a gzipped direct-to-disk body that changes
at constant length (without the pre-sample before the decoded temp is renamed,
it reads as unchanged), and a page with a fixed payload behind a redirect whose
target is renamed, so its file on disk changes length while its bytes do not.
Both were checked against builds with the respective fix reverted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>

* changes: document the renamed-file case in the format page

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>

* changes: refresh two stale test comments

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>

* Merge origin/master into feat/change-report

Both sides appended to tests/Makefile.am's TESTS; kept master's
86_local-proxytrack-cache-longfields.test alongside 88_local-changes.test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>

* changes: translate the new GUI strings into the remaining 28 locales (#740)

The feature PR added the two LANG_CHANGES* entries to lang.def with English and
Francais only; every other locale fell back to English in the WebHTTrack form.
Each file is written in its own declared charset.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* changes: fix the review's blocking findings

Lock the report path against the FTP thread the crawl never joins, seal the
accumulator once the report is written, key entries off the project directory
so the report survives --cache=0, stop calling a file gone when the crawl only
failed to re-fetch it, and skip the hook's work entirely when --changes is off.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* changes: prove the fixes, and give the report a cache-off mode

Adds a changes-race self-test (the FTP shape: notifier threads against the
report path), fixtures for a transfer the crawl never completes, for a leftover
file at a name the crawl mirrors fresh, and for a cache-off mirror, plus a pass
with purging on. Registers the web GUI's --changes box with the clearing
companion master's #725 now requires, and documents the degraded mode.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* changes: keep the failed-transfer case portable

A connection killed before the status line surfaces differently on macOS, where
it truncates the mirrored file to zero (#748). Assert only what holds on both:
it is never reported gone, and its file survives.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <xroche@gmail.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 06:58:33 +02:00
Xavier Roche
d3d3bce8af cmdguide: add a WARC/WACZ archive recipe (#680)
The command-line guide's recipe section had no entry for WARC output, so
the five --warc* flags were discoverable only through --help or the man
page. Add a §11 recipe leading with the gotcha that trips people: the
archive is written alongside the browsable mirror, not instead of it.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 08:09:16 +02:00
Xavier Roche
febdd08cae Named footer fields for -%F (#667)
* Add named {addr}{path}{date}{version} footer fields alongside the legacy %s form

The -%F footer was a bare positional printf: each %s consumed the next of a
fixed addr/path/date/version arg list. It is order-coupled and inexpressive
(you cannot reach {date} without emitting addr and path first, cannot reorder,
and a wrong %s count silently yields "???" or shifted values), and the help
text advertised a cryptic bracket syntax.

hts_footer_format now dispatches on content: a footer containing %s keeps the
legacy positional model byte-for-byte (the default footer and every existing
-%F string are unaffected), while a footer without %s uses named fields
{addr} {path} {date} {version}, with "{{"/"}}" for a literal brace and any
unrecognized {...} left verbatim so typos stay visible. Sanitization is
unchanged: addr/path still pass through html_inline_safe() at the call site,
so the comment-injection guard (#165) still holds.

Driven by a new -#test=footerfmt self-test (tests/01_engine-footerfmt.test)
covering both models, brace escaping, the mixed-mode dispatch boundary, and
the overflow/zero-size return paths.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* Regenerate html/httrack.man.html for the -%F help change

The man/html sync CI guard requires html/httrack.man.html to change whenever
man/httrack.1 does; the -%F help-text update regenerated the man page but not
its html rendering.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* Add {url}{lastmodified}{mime}{charset}{status}{size} footer fields

Extends the named footer set beyond addr/path/date/version. hts_footer_format
now takes a {name,value} table instead of four fixed parameters, so the field
set is data rather than signature and the legacy %s path looks addr/path/date/
version up by name (order-independent, still byte-for-byte). The call site
supplies the new fields from the current response: {url} (scheme + host + path,
credentials stripped like {addr}), {lastmodified}, {mime}, {charset}, {status}
and {size}.

Every network-derived string is html_inline_safe()'d, since the footer sits
inside an HTML comment and a value holding "-->" would otherwise close it and
inject markup (#165); {status}/{size} are formatted integers and need none.
Documented in --help, the man pages and the Command-Line Guide, and covered by
the -#test=footerfmt self-test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 18:15:12 +02:00
Xavier Roche
d2e94b1c99 Fix the download-PDFs example in the command-line guide (#658)
* Fix the download-PDFs example in the command-line guide

The old example ran httrack with no filters, so it mirrored the whole
site rather than downloading its PDFs. The corrected command restricts
to the site, admits the HTML pages as scaffolding for link discovery
(including directory-index pages via *[file]/ so sub-directory PDFs are
found), keeps the PDFs, and drops everything else.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* Use *[path]/ so the PDF recipe follows nested directory indexes

*[file] cannot cross a slash, so *[file]/ only admits a single directory level and
silently drops PDFs linked under nested paths (dated archives, doc trees). *[path]/
spans slashes and follows directory-index pages at any depth, with no over-admission
(only trailing-slash URLs match, never files).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 12:01:02 +02:00
Xavier Roche
2b5e740b55 Document the filter wildcards in the command-line guide (#659)
Add a "Filter wildcards" subsection listing the bracket wildcard forms
the matcher (src/htsfilters.c) implements: *, *[file]/*[name], *[path],
*[param], character sets and ranges, literal escapes, size rules, and
the *[] end anchor. It links to the full filters page for the rest.
Added as an h4 under the filters section so section numbering and the
page ToC are untouched.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 09:59:50 +02:00
Xavier Roche
71c1764525 Advertise -%N's long option and document the real long forms (#652)
* Advertise -%N's long option and document the real long forms

The help display appends each option's long alias by looking it up in the
alias table, but it first stripped a trailing N to turn placeholders like
cN into c. That also turned -%N into -%, so --delayed-type-check was never
advertised in --help (nor in the generated man page). Try the flag as-is
first and only strip a trailing N on a miss; -%N now shows its long form,
and cN still resolves to --sockets.

The command-line guide had several rows marked short-only that in fact
have long aliases (I had read them off a stale installed binary): fill in
--pause, --strip-query, --disable-compression, the three --keep-* dedup
opts, and --delayed-type-check. Regenerate man/httrack.1 to match, and
drop the long-dead commented-out -%O chroot help line.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* Bound the option-token sscanf to the buffer size

CodeQL flagged the %s read into cmd[32] as an unbounded copy on the line
the previous commit touched. The tokens come from HTTrack's own help
strings, not hostile input, so it was not reachable in practice, but the
copy should be bounded regardless. Limit it to %30s (the buffer holds the
leading '-' plus 30 chars and a NUL).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 09:21:50 +02:00
Xavier Roche
db31be5a6c Add a task-oriented command-line guide to the offline docs (#649)
* Add a task-oriented command-line guide to the offline docs

The command-line docs so far are the option list and the generated manual
page: exhaustive, but organized by flag, not by task. Newcomers arrive
expecting wget/curl syntax and hit the same walls (only the index came
down, filters that do the opposite of what they read, an update that
deletes files), because the reference answers "what does -X do", not "how
do I do the thing I want".

cmdguide.html is a task-oriented layer on top of the manual page: quick
start, scope, filters, limits, naming, links, identity/login, proxy,
update/cache, scripting, then eleven copy-ready recipes with the one
gotcha each. It foregrounds the defaults that actually surprise people
(the ~100 KB/s rate cap, the -c/-A/-%c security clamps, --update purge
semantics, the depth off-by-one, mime filters running after headers) and
links the manual page for per-option detail. Linked from cmddoc.html and
the index. Every documented default and behavior was checked against the
engine source.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* Correct four scope/option descriptions in the guide

Adversarial review against the engine source caught four mislabels:
-d is "same principal domain" and -l is "same TLD" (not "same directory"
and "same domain"); -U goes up only and -B goes both ways (the guide had
-U as up-and-down and -B as "anywhere"); -%M archives the whole mirror
into one index.mht, not each page; and -t HEAD-tests links outside scope
rather than reporting "what a scope would reach". The rest of the guide's
documented behavior verified against code.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* Fix the PDF recipe, prefer long options, drop the section rules

Review feedback from Xavier:

- The "grab every PDF" recipe was wrong. `-* +*.pdf` blocks the HTML pages
  that carry the PDF links, so the crawl only keeps PDFs linked from the
  entry page (verified on a local fixture). There is no PDF-only crawl:
  HTTrack finds PDFs by parsing HTML. Reworded to let the site traverse,
  with the off-host case handled by a `+host/*.pdf` rule.
- Recipes now use long options throughout. `-%!` in particular is replaced
  by `--disable-security-limits`: the bare `!` triggers shell history
  expansion and is easy to fumble, and the long name says what it does.
- Dropped the per-section `<hr>` rules; the sibling doc pages don't use
  them and the section headings already separate the content.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* Key the reference tables by long option name, short in parens

Follows the recipe conversion: every table row now leads with the long
option (--depth, --stay-on-same-domain, ...) and carries the short flag in
parentheses. The guru %-flags that have no long form (-%g, -%j, -%o, -%y,
-%z, -%G, -%t, -%N) stay short. Also switches the remaining inline command
snippets in descriptions to long form (--assume, --structure,
--user-agent "").

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 07:27:04 +02:00