Compare commits

...

24 Commits

Author SHA1 Message Date
Xavier Roche
30a4d37b0e sitemap: keep the ingestion state out of htsoptstate
htsoptstate is embedded by value as httrackp.state, so a field at its tail
shifts every httrackp member declared after it: an offsetof probe put
warc_file at 141752 on master and 141760 on the branch. Move the pointer to
httrackp's own tail, where every existing offset holds and copy_htsopt still
ignores it.

Also renumber the crawl test to 89, master having taken 87 and 90, and give
the new option8 checkbox the hidden companion that 90_webhttrack-checkbox-clear
requires, plus its row in that test's table.

Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-26 22:18:40 +02:00
Xavier Roche
97f9a04d31 Merge remote-tracking branch 'origin/master' into feat/sitemap
Signed-off-by: Xavier Roche <roche@httrack.com>

# Conflicts:
#	src/htsselftest.c
#	tests/Makefile.am
2026-07-26 22:10:56 +02:00
Xavier Roche
ce7dcfa9de Options ticked on by default cannot be un-ticked in the web GUI (#725)
* Options ticked on by default cannot be un-ticked in the web GUI

An unchecked HTML checkbox posts nothing, so htsserver never overwrites the
value it already holds. Every box in the wizard is one-way: once the stored
value is "1", whether htsserver seeded it at startup, a loaded profile set it,
or the user ticked it earlier in the session, un-ticking and submitting leaves
the option on and draws the box ticked again. Only four boxes, all in
option1.html, carried the companion hidden field that guards against this.

Add it to every remaining bare checkbox, and switch cookies and parsejava to
${ztest:...} so a cleared box emits --cookies=0 / --parse-java=0; ${test:...}
renders nothing at all when the value is empty, which is not "off" for an
option the engine turns on by default.

index, urlhack and keep-alive are deliberately left alone: their long options
are declared "single" in htsalias.c and optalias_check drops the =value, so
--index=0 resolves to -I and turns the option back on. That parser bug is
pre-existing and needs its own fix.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: pin last-write-wins and cover every checkbox

The runtime leg posted each name once, so a first-wins body parser would
have passed while the fix did nothing in a real browser: post the
duplicated name in both orders and assert the last value wins. Replace the
four hand-written option cases with a table covering all 28 non-skipped
boxes, asserting the command-line token and the Windows-profile key each
state emits, plus a completeness check so a new box cannot slip through
unexercised.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: rename the loop variable shadowing the scanned page

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 22:08:46 +02:00
Xavier Roche
069573edc3 Test assertions read a padded or truncated reply as a clean security verdict (#728)
* tests: a failed request must not read as a clean security verdict

Under pipefail, "request | grep -q MARKER && fail" skips the fail when the
request itself errors: the leak checks in tests 78 and 85 then pass without
ever having run. Capture the reply first and fail loudly if it never arrived.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* AGENTS.md: record the fail-open assertion shape

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: a reply that proves nothing must not read as a clean verdict

The previous commit converted some of the fail-open assertions and left three.
78's refusal loop still piped into "grep -q ... && fail": grep -q exits on the
first match and SIGPIPEs the producer, so under pipefail a hostile reply that
pads its Location past the 64 KB pipe buffer suppresses the failure exactly as
a dead probe would. 85's fetch() only required a non-empty reply, so a
truncated body or a 302 to the file passed the leak checks marker-free, and no
assertion looked at the status line at all. 78's store probe had no emptiness
guard, so an empty page read as "the store was not written".

Match from here-strings throughout, give fetch() the status each caller
expects, and route 78's store probe through a helper that requires a served
page. 77's X-Injected check had the same shape.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 22:08:10 +02:00
Xavier Roche
a75f437df9 lienrelatif() reads one byte before its stack buffer on an empty path (#729)
The trim that walks back to the last '/' starts at `curr + strlen(curr) - 1`,
which is `curr - 1` when the path is empty. The loop then dereferences it.

An empty path is reachable today: the pre-pass that strips a query does
`strncatbuff(newcurr_fil, curr_fil, a - curr_fil)`, so any `curr_fil` starting
with '?' hands the walk an empty string. `-#test=relative "dir/page.html" "?x"`
under ASan reports the underflow.

The read is one byte and the loop stops immediately either way, so the guard
changes no output: over the 484 ordered pairs of a 22-value path corpus, run
against builds that force the byte before the buffer to 0 and to '/', the
guarded and unguarded results are identical.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 22:06:11 +02:00
Xavier Roche
783f6ee1f5 AGENTS.md: record what the msg[80] hardening batch taught (#737)
Five PRs across the engine and ProxyTrack turned up the same few traps
more than once, and none of them are obvious from the code.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 21:39:57 +02:00
Xavier Roche
1f944c9ef7 sitemap: date the new files 2026
The headers were copied from an existing file and kept its 1998 year.

Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-26 21:21:54 +02:00
Xavier Roche
913caf68be Clear the last three compiler warnings (#733)
* Clear the last three compiler warnings

finalurl was sized for one of the two URLs it concatenates. The IIS-bug
example callback overwrites a suffix in place with a same-length
replacement and must not terminate the string, which is memcpy, not
strncpy. The coucal bench's if/else chain has no final else, so result was
only initialised on the paths gcc could not prove exhaustive.

A clean build now reports zero warnings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* Bound the IIS suffix copy by what matched, not by the table

Copying strlen(replacement) leaves the "MUST be the same sizes" comment as
the only thing standing between a future table edit and an overflow. j is
the number of bytes just matched in the destination, so copying j is safe
whatever the table holds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 20:37:52 +02:00
Xavier Roche
59660102d6 A cache field wider than ours aborts the engine instead of clipping (#732)
* A cache field wider than ours aborts the engine instead of clipping

The read-side ZIP_READFIELD_STRING used strlcpybuff, and the whole *_safe_
family aborts on overflow rather than truncating. Since the header line is
bounded only by HTS_URLMAXSIZE and msg is 80 bytes, a cache written by
another build, or a corrupt one, kills the crawl outright. The corrupt-cache
self-test already promises "rejected per-entry, never crash".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* Pin the clip to each field's own capacity

Review found one case exercised only msg[80], so a hardcoded clip length
passed. lastmodified[64] is narrower, and no single constant satisfies
both. The forged replacement was also one byte longer than the line it
overwrote, which only worked because corrupt_patch copies exactly the
pattern length.

Also stop claiming another build's cache can trigger this: the writer
emits each field from the same struct the reader fills, so it takes a
corrupt cache.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 20:37:38 +02:00
Xavier Roche
d22dae4895 Merge remote-tracking branch 'origin/master' into feat/sitemap
Signed-off-by: Xavier Roche <roche@httrack.com>

# Conflicts:
#	tests/Makefile.am
2026-07-26 20:25:38 +02:00
Xavier Roche
ec96d5f24a selftest: bound the sitemap document builders' snprintf accumulation
snprintf returns the length it wanted to write, so accumulating it blind
lets the next offset and size argument walk past the buffer. Guard each
step the way the argv builder above already does, and give the per-URL
loop a real remaining-space bound instead of a fixed 33.

Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-26 20:21:11 +02:00
Xavier Roche
e96399910b Share the in-progress display struct instead of copying it (#734)
t_StatsBuffer was defined twice, byte for byte, in httrack.h and htsweb.h,
with NStatsBuffer duplicated alongside. httrack and htsserver each keep
their own array, so nothing catches the two drifting apart, and the last
change to the struct had to be applied to both by hand. Both now include
src/htsstats.h; sizes and offsets are unchanged.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 20:19:01 +02:00
Xavier Roche
52d0ab2356 An entry with no usable Last-Modified crashes proxytrack --convert (#731)
* Cached entry with no usable date crashes the ARC writer

PT_SaveCache__Arc_Fun dereferenced convert_time_rfc822() straight into the
record line, so any entry whose Last-Modified is absent or unparseable took
proxytrack --convert down. The sibling caller a thousand lines up already
guards the same call; this one fills in the epoch instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* Assert the archive date, not just the entry

Review found the test blind to the two mutants that matter: a guard firing
unconditionally clobbers every valid date to the epoch, and one that skips
the year and day emits a month and day of 00. Grepping only for the URL saw
neither. Assert the date field exactly, and add a valid-date case so the
untouched path is pinned too.

A bare "Last-Modified: 0" crashes the same way, so it joins the cases.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 20:18:58 +02:00
Xavier Roche
4f15490186 fuzz: keep only the four sitemap seed inputs
A libFuzzer run writes its finds into the first corpus directory, and 191 of
them were committed with the harness.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-26 19:22:42 +02:00
Xavier Roche
028ac8b5ad sitemap: add the fuzz harness the parser was missing, and drop truncated Sitemap: lines
The file header called the scanner fuzzable while fuzz/ registered ten
harnesses and none for it. fuzz-sitemap feeds it raw XML, gzip-framed bodies
and truncated streams off a heap copy of exactly the input size, so an overread
is an ASan report rather than a quiet pass, with a four-file seed corpus.
60000 runs clean under ASan+UBSan.

robots_parse now drops a Sitemap: line that filled its scratch buffer instead
of handing on the half URL it was truncated to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-26 19:21:06 +02:00
Xavier Roche
3745d8321b sitemap: gate each fetch by who asked for it, and anchor travel on the start URL
A nine-agent review found a scope escape: the sitemap document was its own
`premier`, so the wizard measured travel from wherever the site chose to put
its sitemap. A root /sitemap.xml therefore widened a /deep/dir/ crawl to the
whole host. The ingester now points the wizard at the crawl's own start link
and lets each seeded URL become its own anchor, which is what a command-line
seed gets.

Robots handling was both mistimed and undifferentiated. The Sitemap: lines are
now collected by robots_parse, on the same body in the same fetch, and acted on
after the parsed rules are installed rather than before; and the decision comes
from a new hts_robots_forbids extracted out of the wizard, so the sitemap path
inherits the -s1 filters-win override instead of a stricter hand-rolled check.
The four fetches are no longer treated alike: a sitemap the user names is user
intent, one the site declares invites the fetch, only the guessed /sitemap.xml
obeys a Disallow, and the URLs listed inside stay fully gated.

Also: hts_unescapeEntities replaces the private entity decoder, whose guard
tests and fuzz corpus it silently forfeited; hts_codec_head replaces hts_zhead,
which is only defined under HTS_USEZLIB; the composed URL buffer now fits two
maximal components plus a scheme, which a 2046-byte --sitemap-url reached; the
bounded search is promoted to htstools as hts_memstr; and the live state moves
from httrackp into htsoptstate, leaving two installed fields rather than three.

Tests gain the scope escape, the three robots cases, a cap-boundary control,
the handler invocation count and a compression-bomb decode. Every one was
checked against a deliberately broken build.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-26 19:10:01 +02:00
Xavier Roche
eea8ec5b29 Merge remote-tracking branch 'origin/master' into feat/sitemap
# Conflicts:
#	tests/Makefile.am
2026-07-26 18:55:22 +02:00
Xavier Roche
0ce7da1973 sitemap: a child sitemap is a fetch, so filters and robots.txt must gate it
An adversarial review found that a <sitemapindex> <loc>, and a robots.txt
Sitemap: line, went straight to hts_record_link: the request went out even
when a -* rule or robots.txt Disallow covered it. Only the <urlset> half ran
through the wizard. Gate the document itself on the filters and on
robots.txt, which is all that can apply: the wizard proper wants a referring
link, and its up/down travel rules would judge a child sitemap against the
parent sitemap's own directory. The robots.txt probe is exempt, being the
request that fetches the rules.

A 301 also used to end ingestion silently, since the engine re-queues the
target as a fresh link that carried no sitemap marking. That hit any site
redirecting http to https. The marking now follows the redirect.

The "N URL(s) added" counter reported what the scanner emitted rather than
what was taken, which hid both of the above; it now reads "N of M". The
fallback to /sitemap.xml keys on the same corrected count, so a robots.txt
whose only Sitemap: line is off-host or filtered still falls back. Root
classification skips a UTF-8 BOM and an XML namespace prefix, and the doc
list is cleared when a mirror starts rather than only when it ends.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-26 18:14:41 +02:00
Xavier Roche
dcfc4acef8 sitemap: state what the decompression cap actually binds on
deflate tops out near 1032:1, so hts_codec_maxout never binds before the
64 MiB cap; the old comment implied a ratio guard that cannot fire.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-26 18:00:57 +02:00
Xavier Roche
1d96350564 sitemap: fix an out-of-bounds read, tighten the parser and the tests
lienrelatif() walked back from the last character of its current-path
argument without checking the path was non-empty, reading one byte before
the stack buffer. htsAddLink is the first caller to pass an empty savename,
which sitemap documents have because they are ingested rather than mirrored,
so ASan caught it on the new crawl test.

The parser drops a value whose numeric character reference decodes outside
printable ASCII, rather than leaving the reference verbatim and seeding a URL
the site never published, and classifies a document by its real root element,
so a comment naming the other one no longer flips urlset and sitemapindex.
The robots.txt line reader is bounded by the body size instead of relying on
a NUL terminator.

The self-test moves to 01_zlib-sitemap.test: MSan runs 01_engine-* only,
because an uninstrumented libz floods it with false positives.

Tests gain the assertions the earlier ones were missing: which of the
robots.txt route and the /sitemap.xml fallback was taken, that the sitemap
documents stay out of the mirror, that the off-host child sitemap is refused,
the sitemapindex nesting cap, the per-document URL cap at its production
value, and copy_htsopt coverage for the two new fields.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-26 17:59:19 +02:00
Xavier Roche
f2abda8c0e Merge remote-tracking branch 'origin/master' into feat/sitemap 2026-07-26 17:50:13 +02:00
Xavier Roche
bb3d8db103 sitemap: distinguish the sitemapindex log line from a urlset one
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-26 17:50:13 +02:00
Xavier Roche
47fe9558da Merge remote-tracking branch 'origin/master' into feat/sitemap
# Conflicts:
#	tests/Makefile.am
2026-07-26 17:40:46 +02:00
Xavier Roche
9cc9a36fa6 Read sitemap files so URLs nothing links to are found
HTTrack finds URLs only by parsing links, so anything a site publishes solely
in its sitemap stayed invisible: robots.txt was already parsed, but its
Sitemap: lines were ignored and nothing else in the tree touched sitemaps.

Adds opt-in --sitemap (-%m), which probes the start host's robots.txt and
falls back to /sitemap.xml, and --sitemap-url (-%mu) for an explicit document.
Handles <urlset> and nested <sitemapindex>, plain or gzipped. Discovered URLs
enter with the full depth budget but still go through the wizard, so filters
and scope rules decide; a sitemap is not a filter bypass.

The parser reads attacker-controlled XML off the network, so it is capped on
URL count, index nesting, decompressed size and decompression ratio, and child
sitemaps must stay on the host that named them.

Closes #712

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
2026-07-26 17:38:02 +02:00
56 changed files with 2233 additions and 127 deletions

View File

@@ -26,6 +26,12 @@ the operational checklist: toolchain, invariants, and how to ship a change.
check`, or `PATH="<bld>/src:$PATH"` for a manual run.
- Give new `.test` scripts `set -e`: the older ones predate the rule, so several
`local-crawl.sh` calls with no `set -e` report PASS on any non-last failure.
- Never assert with `cmd | grep -q MARKER && fail`. Under `pipefail` the
pipeline is non-zero both when `cmd` fails and when `grep -q` matches early
and SIGPIPEs it, so the `&&` never fires and a probe that proved nothing reads
as "marker absent". Capture the reply, assert the status line it must carry
(an empty, truncated or redirected one is marker-free too), then match with a
here-string.
## Hard invariants
- **Generated autotools files are NOT in git.** `configure`, every
@@ -48,6 +54,16 @@ the operational checklist: toolchain, invariants, and how to ship a change.
- Bounds-check every copy. Overflow-safe form: put the untrusted value alone,
`untrusted < limit - controlled` — never `controlled + untrusted < limit`,
which can wrap and pass.
- **Abort or clip is a decision, not a default.** The `*_safe_` helpers
(`strcpybuff`, `strlcpybuff`, `strcatbuff`) **abort** on overflow. Right for
our own data, wrong for anything read back from a cache, a header or the
wire, where it trades a memory smash for a crash on malformed input. Clip
with `dst[0] = '\0'; strlncatbuff(dst, src, size, size - 1)`.
- **A warning class is not the unsafe set.** `-Wformat-truncation` fires only on
a *bounded* `snprintf` whose return is discarded, so an unbounded `sprintf`
into the same buffer never appears on it. Before scoping a hardening pass off
compiler output, grep the unguarded forms yourself (`\bsprintf\s*\(`,
`\bstrcpy\s*\(`, `\bstrcat\s*\(`).
## C conventions
- **Use the `*t` allocator wrappers, never raw libc** (`htssafe.h`):
@@ -84,6 +100,17 @@ Before pushing, and when reviewing others, don't skim for bugs:
layout/ABI, cache/wire format, or a security path? A static or unit check
isn't enough; exercise the wrong behavior at runtime. Claude Code:
`/review-recipe`.
- **Poison a canary, never compare it against zero.** Checking that a
neighbouring field is still `'\0'` cannot see the stray NUL an off-by-one
terminator writes — the exact bug the canary is there for. Fill it with a
non-zero byte, and prove it by killing both the stray-`'X'` and the
stray-NUL mutant. Neither ASan nor `_FORTIFY_SOURCE` sees an overflow that
lands inside the same struct.
- **Overshoot every destination, not one.** A bounds test that oversizes a
single field cannot tell a per-field bound from a one-size-fits-all one, nor
from a fix that bounds that field and leaves its neighbours raw. Exercise
each destination the path touches, spanning at least two capacities, and
check what the code actually emits before writing the expected values.
## Commits
- **Sign-off is mandatory.** Every commit carries a `Signed-off-by` trailer:

View File

@@ -2,7 +2,7 @@
if FUZZERS
noinst_PROGRAMS = fuzz-charset fuzz-meta fuzz-idna fuzz-entities \
fuzz-unescape fuzz-filters fuzz-url fuzz-header fuzz-cachendx \
fuzz-htsparse
fuzz-htsparse fuzz-sitemap
endif
AM_CPPFLAGS = \
@@ -27,6 +27,7 @@ fuzz_url_SOURCES = fuzz-url.c fuzz.h
fuzz_header_SOURCES = fuzz-header.c fuzz.h
fuzz_cachendx_SOURCES = fuzz-cachendx.c fuzz.h
fuzz_htsparse_SOURCES = fuzz-htsparse.c fuzz.h
fuzz_sitemap_SOURCES = fuzz-sitemap.c fuzz.h
# List corpus files explicitly: automake does not expand EXTRA_DIST globs.
EXTRA_DIST = README.md run-fuzzers.sh \
@@ -47,4 +48,6 @@ EXTRA_DIST = README.md run-fuzzers.sh \
corpus/cachendx/regress-overadvance.bin \
corpus/cachendx/regress-truncated-entry.bin \
corpus/htsparse/basic.html corpus/htsparse/script-inscript.html \
corpus/htsparse/meta-usemap.html corpus/htsparse/malformed.html
corpus/htsparse/meta-usemap.html corpus/htsparse/malformed.html \
corpus/sitemap/urlset.xml corpus/sitemap/sitemapindex.xml \
corpus/sitemap/truncated.xml corpus/sitemap/urlset.xml.gz

View File

@@ -0,0 +1 @@
<sitemapindex><sitemap><loc>http://h.test/s2.xml.gz</loc></sitemap></sitemapindex>

View File

@@ -0,0 +1 @@
<urlset><loc>http://h.test/x

View File

@@ -0,0 +1 @@
<?xml version="1.0"?><urlset><url><loc>http://h.test/a.html</loc></url><url><loc>https://h.test/b?x=1&amp;y=2</loc></url></urlset>

Binary file not shown.

60
fuzz/fuzz-sitemap.c Normal file
View File

@@ -0,0 +1,60 @@
/* ------------------------------------------------------------ */
/*
HTTrack Website Copier, Offline Browser for Windows and Unix
Copyright (C) 2026 Xavier Roche and other contributors
SPDX-License-Identifier: GPL-3.0-or-later
This program is free software: you can redistribute it and/or modify
it under the terms of the GNU General Public License as published by
the Free Software Foundation, either version 3 of the License, or
(at your option) any later version.
This program is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
GNU General Public License for more details.
You should have received a copy of the GNU General Public License
along with this program. If not, see <http://www.gnu.org/licenses/>.
Ethical use: we kindly ask that you NOT use this software to harvest email
addresses or to collect any other private information about people. Doing so
would dishonor our work and waste the many hours we have spent on it.
Please visit our Website: http://www.httrack.com
*/
/* Fuzz the sitemap <loc> scanner (htssitemap.c): raw XML, gzip-framed bodies
and truncated streams all arrive here straight off the network. */
#include "fuzz.h"
#include "htssitemap.h"
static hts_boolean sm_count(void *arg, const char *url) {
int *const n = (int *) arg;
(void) url;
(*n)++;
return HTS_TRUE;
}
int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size) {
static const int caps[] = {0, 1, 16, HTS_SITEMAP_MAX_URLS_DOC};
hts_boolean is_index;
char *body;
int n = 0, cap;
if (size == 0)
return 0;
cap = caps[data[0] % (sizeof(caps) / sizeof(caps[0]))];
data++, size--;
/* A heap copy of exactly `size` bytes: the scanner must never rely on a
terminator, and ASan turns any overread into a report. */
body = malloct(size != 0 ? size : 1);
memcpy(body, data, size);
(void) hts_sitemap_scan(body, size, cap, &is_index, sm_count, &n);
freet(body);
return 0;
}

View File

@@ -163,8 +163,26 @@ the index" problems disappear.</p>
<tr><td><tt>--near (-n)</tt></td><td>Also fetch non-HTML files "near" a followed link, such as an image linked from a page you kept but hosted elsewhere.</td></tr>
<tr><td><tt>--ext-depth (-%e)</tt></td><td>How many levels of external links to follow once the crawl leaves your scope (default 0).</td></tr>
<tr><td><tt>--test (-t)</tt></td><td>Also HEAD-test links that fall outside the scope, which are normally refused, without downloading them: a way to see what scope is excluding.</td></tr>
<tr><td><tt>--sitemap (-%m), --sitemap-url URL (-%mu)</tt></td><td>Also take start URLs from the site's sitemap, for pages nothing links to. Off by default.</td></tr>
</table>
<p>Link-following only finds what something links to. Anything a site publishes
solely in its sitemap is invisible to HTTrack unless you ask for it.
<tt>--sitemap</tt> reads the start host's <tt>robots.txt</tt> for
<tt>Sitemap:</tt> lines and falls back to <tt>/sitemap.xml</tt>;
<tt>--sitemap-url</tt> names one directly. Nested <tt>sitemapindex</tt> files
and gzipped <tt>.xml.gz</tt> sitemaps are followed. The URLs found become start
URLs with the full depth budget, but they still go through your filters and
scope rules, so a sitemap cannot widen a crawl you deliberately narrowed. It is
off by default because a sitemap can list thousands of pages nothing links
to.</p>
<p>One surprise worth knowing: a sitemap you name with <tt>--sitemap-url</tt>,
and one the site itself declares in <tt>robots.txt</tt>, are fetched even when
<tt>robots.txt</tt> disallows that path, because naming or declaring a sitemap
is an invitation to read it. Only the guessed <tt>/sitemap.xml</tt> obeys a
<tt>Disallow</tt>. The URLs listed inside are gated normally either way.</p>
<p>The single most common surprise is "only the home page came down." That is
usually not a scope option at all: it is an off-host redirect. A start URL of
<tt>http://example.com/</tt> that redirects to <tt>https://www.example.com/</tt>

View File

@@ -87,8 +87,8 @@ offline browser : copy websites to a local directory</p>
--host-control[=N]</b> ] [ <b>-%P,
--extended-parsing[=N]</b> ] [ <b>-n, --near</b> ] [ <b>-t,
--test</b> ] [ <b>-%L, --list</b> ] [ <b>-%S, --urllist</b>
] [ <b>-NN, --structure[=N]</b> ] [ <b>-%N,
--delayed-type-check</b> ] [ <b>-%D,
] [ <b>-%m, --sitemap</b> ] [ <b>-NN, --structure[=N]</b> ]
[ <b>-%N, --delayed-type-check</b> ] [ <b>-%D,
--cached-delayed-type-check</b> ] [ <b>-%M, --mime-html</b>
] [ <b>-LN, --long-names[=N]</b> ] [ <b>-KN,
--keep-links[=N]</b> ] [ <b>-x, --replace-external</b> ] [
@@ -575,6 +575,22 @@ URL per line) (--list &lt;param&gt;)</p></td></tr>
<p>&lt;file&gt; add all scan rules located in this text
file (one scan rule per line) (--urllist &lt;param&gt;)</p></td></tr>
<tr valign="top" align="left">
<td width="9%"></td>
<td width="4%">
<p>-%m</p></td>
<td width="5%"></td>
<td width="82%">
<p>seed the crawl from the site&rsquo;s sitemap (robots.txt
Sitemap:, then /sitemap.xml); --sitemap-url URL names one
explicitly. A sitemap you name, or one the site declares, is
fetched even under robots.txt Disallow; only the guessed
/sitemap.xml obeys it. The URLs found still pass every
filter and scope rule (--sitemap)</p></td></tr>
</table>
<h3>Build options:

View File

@@ -98,6 +98,9 @@ ${do:end-if}
<input type="hidden" name="redirect" value="">
<input type="hidden" name="closeme" value="">
<!-- clear if not checked -->
<input type="hidden" name="ftpprox" value="">
${LANG_PROXYTYPE}
<select name="proxytype"
title='${html:LANG_PROXYTYPETIP}' onMouseOver="info('${html:LANG_PROXYTYPETIP}'); return true" onMouseOut="info('&nbsp;'); return true"

View File

@@ -98,6 +98,13 @@ ${do:end-if}
<input type="hidden" name="redirect" value="">
<input type="hidden" name="closeme" value="">
<!-- clear if not checked -->
<input type="hidden" name="errpage" value="">
<input type="hidden" name="external" value="">
<input type="hidden" name="hidepwd" value="">
<input type="hidden" name="hidequery" value="">
<input type="hidden" name="nopurge" value="">
${LANG_I33}
<br>
<select name="build"

View File

@@ -98,6 +98,9 @@ ${do:end-if}
<input type="hidden" name="redirect" value="">
<input type="hidden" name="closeme" value="">
<!-- clear if not checked -->
<input type="hidden" name="windebug" value="">
${LANG_I40c}
<br>

View File

@@ -98,6 +98,11 @@ ${do:end-if}
<input type="hidden" name="redirect" value="">
<input type="hidden" name="closeme" value="">
<!-- clear if not checked -->
<input type="hidden" name="ka" value="">
<input type="hidden" name="remt" value="">
<input type="hidden" name="rems" value="">
<table border="0" width="100%" cellspacing="0">
<tr><td>

View File

@@ -98,6 +98,18 @@ ${do:end-if}
<input type="hidden" name="redirect" value="">
<input type="hidden" name="closeme" value="">
<!-- clear if not checked -->
<input type="hidden" name="cookies" value="">
<input type="hidden" name="parsejava" value="">
<input type="hidden" name="updhack" value="">
<input type="hidden" name="urlhack" value="">
<input type="hidden" name="keepwww" value="">
<input type="hidden" name="keepslashes" value="">
<input type="hidden" name="keepqueryorder" value="">
<input type="hidden" name="toler" value="">
<input type="hidden" name="http10" value="">
<input type="hidden" name="sitemap" value="">
<input type="checkbox" name="cookies" ${checked:cookies}
title='${html:LANG_I1b}' onMouseOver="info('${html:LANG_I1b}'); return true" onMouseOut="info('&nbsp;'); return true"
> ${LANG_I58}
@@ -132,6 +144,17 @@ ${listid:robots:LISTDEF_8}
</select>
<br><br>
<input type="checkbox" name="sitemap" ${checked:sitemap}
title='${html:LANG_SITEMAPTIP}' onMouseOver="info('${html:LANG_SITEMAPTIP}'); return true" onMouseOut="info('&nbsp;'); return true"
> ${LANG_SITEMAP}
<br><br>
${LANG_SITEMAPURL}
<input name="sitemapurl" value="${sitemapurl}" size="40"
title='${html:LANG_SITEMAPURLTIP}' onMouseOver="info('${html:LANG_SITEMAPURLTIP}'); return true" onMouseOut="info('&nbsp;'); return true"
>
<br><br>
<input type="checkbox" name="updhack" ${checked:updhack}
title='${html:LANG_I1k}' onMouseOver="info('${html:LANG_I1k}'); return true" onMouseOut="info('&nbsp;'); return true"
> ${LANG_I62b}

View File

@@ -98,6 +98,13 @@ ${do:end-if}
<input type="hidden" name="redirect" value="">
<input type="hidden" name="closeme" value="">
<!-- clear if not checked -->
<input type="hidden" name="warc" value="">
<input type="hidden" name="norecatch" value="">
<input type="hidden" name="logf" value="">
<input type="hidden" name="index" value="">
<input type="hidden" name="index2" value="">
<input type="checkbox" name="cache2" ${checked:cache2}
title='${html:LANG_I1e}' onMouseOver="info('${html:LANG_I1e}'); return true" onMouseOut="info('&nbsp;'); return true"
> ${LANG_I61}

View File

@@ -141,6 +141,8 @@ ${do:copy:KeepSlashes:keepslashes}
${do:copy:KeepQueryOrder:keepqueryorder}
${do:copy:StripQuery:stripquery}
${do:copy:StoreAllInCache:cache2}
${do:copy:Sitemap:sitemap}
${do:copy:SitemapUrl:sitemapurl}
${do:copy:Warc:warc}
${do:copy:WarcFile:warcfile}
${do:copy:LogType:logtype}

View File

@@ -114,7 +114,7 @@ ${do:end-if}
${/* Real commands and ini file generated below */}
<!-- engine commandline -->
<!-- engine commandline; ztest so a cleared default-on option still emits its disabling flag -->
${do:output-mode:html}
<textarea name="command" cols="50" rows="4" style="visibility:hidden">
httrack \
@@ -175,8 +175,8 @@ ${/* -m<n> resets the html limit, so the bare form must precede the -m,<n> one *
\
${unquoted:url2}
\
${test:cookies:--cookies=0:}
${test:parsejava:--parse-java=0:}
${ztest:cookies:--cookies=0:}
${ztest:parsejava:--parse-java=0:}
${test:updhack:--updatehack}
${test:urlhack:--urlhack=0:--urlhack}
${test:keepwww:--keep-www-prefix}
@@ -188,6 +188,8 @@ ${/* -m<n> resets the html limit, so the bare form must precede the -m,<n> one *
${test:toler:--tolerant}
${test:http10:--http-10}
${test:cache2:--store-all-in-cache}
${test:sitemap:--sitemap}
${test:sitemapurl:--sitemap-url "}${html:sitemapurl}${test:sitemapurl:"}
${test:warc:--warc}
${test:warcfile:--warc-file "}${arg:warcfile}${test:warcfile:"}
${test:norecatch:--do-not-recatch}
@@ -240,6 +242,8 @@ KeepSlashes=${ztest:keepslashes:0:1}
KeepQueryOrder=${ztest:keepqueryorder:0:1}
StripQuery=${stripquery}
StoreAllInCache=${ztest:cache2:0:1}
Sitemap=${ztest:sitemap:0:1}
SitemapUrl=${sitemapurl}
Warc=${ztest:warc:0:1}
WarcFile=${warcfile}
LogType=${logtype}

View File

@@ -1042,3 +1042,11 @@ LANG_WARCFILE
WARC archive name:
LANG_WARCFILETIP
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
LANG_SITEMAP
Seed the crawl from the site's sitemap
LANG_SITEMAPTIP
Read the site's sitemap (robots.txt Sitemap: lines, then /sitemap.xml) and add every URL it lists as a start URL.
LANG_SITEMAPURL
Sitemap address:
LANG_SITEMAPURLTIP
Address of a sitemap to read instead of probing the site; leave blank to probe robots.txt then /sitemap.xml.

View File

@@ -1012,3 +1012,11 @@ WARC archive name:
WARC archive name:
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Seed the crawl from the site's sitemap
Seed the crawl from the site's sitemap
Read the site's sitemap (robots.txt Sitemap: lines, then /sitemap.xml) and add every URL it lists as a start URL.
Read the site's sitemap (robots.txt Sitemap: lines, then /sitemap.xml) and add every URL it lists as a start URL.
Sitemap address:
Sitemap address:
Address of a sitemap to read instead of probing the site; leave blank to probe robots.txt then /sitemap.xml.
Address of a sitemap to read instead of probing the site; leave blank to probe robots.txt then /sitemap.xml.

View File

@@ -1012,3 +1012,11 @@ WARC archive name:
Nom de l'archive WARC :
Optional base name for the WARC archive; leave blank to auto-name it under the output directory.
Nom de base optionnel pour l'archive WARC ; laissez vide pour le générer automatiquement dans le répertoire de sortie.
Seed the crawl from the site's sitemap
Partir du plan de site (sitemap)
Read the site's sitemap (robots.txt Sitemap: lines, then /sitemap.xml) and add every URL it lists as a start URL.
Lire le plan de site (lignes Sitemap: de robots.txt, puis /sitemap.xml) et ajouter chaque URL listée comme adresse de départ.
Sitemap address:
Adresse du plan de site :
Address of a sitemap to read instead of probing the site; leave blank to probe robots.txt then /sitemap.xml.
Adresse d'un plan de site à lire au lieu de sonder le site ; laissez vide pour sonder robots.txt puis /sitemap.xml.

View File

@@ -71,7 +71,9 @@ static int mysavename(t_hts_callbackarg * carg, httrackp * opt,
for(j = 0; iisBogus[i][j] == a[j] && iisBogus[i][j] != '\0'; j++) ;
if (iisBogus[i][j] == '\0'
&& (a[j] == '\0' || a[j] == '/' || a[j] == '\\')) {
strncpy(a, iisBogusReplace[i], strlen(iisBogusReplace[i]));
/* j bytes matched, so j fit: copying j cannot overrun whatever the
table holds, and the tail must survive untouched */
memcpy(a, iisBogusReplace[i], (size_t) j);
break;
}
}

View File

@@ -3,7 +3,7 @@
.\"
.\" This file is generated by man/makeman.sh; do not edit by hand.
.\" SPDX-License-Identifier: GPL-3.0-or-later
.TH httrack 1 "23 July 2026" "httrack website copier"
.TH httrack 1 "26 July 2026" "httrack website copier"
.SH NAME
httrack \- offline browser : copy websites to a local directory
.SH SYNOPSIS
@@ -36,6 +36,7 @@ httrack \- offline browser : copy websites to a local directory
[ \fB\-t, \-\-test\fR ]
[ \fB\-%L, \-\-list\fR ]
[ \fB\-%S, \-\-urllist\fR ]
[ \fB\-%m, \-\-sitemap\fR ]
[ \fB\-NN, \-\-structure[=N]\fR ]
[ \fB\-%N, \-\-delayed\-type\-check\fR ]
[ \fB\-%D, \-\-cached\-delayed\-type\-check\fR ]
@@ -187,6 +188,8 @@ test all URLs (even forbidden ones) (\-\-test)
<file> add all URL located in this text file (one URL per line) (\-\-list <param>)
.IP \-%S
<file> add all scan rules located in this text file (one scan rule per line) (\-\-urllist <param>)
.IP \-%m
seed the crawl from the site's sitemap (robots.txt Sitemap:, then /sitemap.xml); \-\-sitemap\-url URL names one explicitly. A sitemap you name, or one the site declares, is fetched even under robots.txt Disallow; only the guessed /sitemap.xml obeys it. The URLs found still pass every filter and scope rule (\-\-sitemap)
.SS Build options:
.IP \-NN
structure type (0 *original structure, 1+: see below) (\-\-structure[=N])

View File

@@ -48,7 +48,7 @@ htsserver_LDFLAGS = $(AM_LDFLAGS) $(LDFLAGS_PIE)
lib_LTLIBRARIES = libhttrack.la
htsserver_SOURCES = htsserver.c htsserver.h htsweb.c htsweb.h \
htsserver_SOURCES = htsserver.c htsserver.h htsweb.c htsweb.h htsstats.h \
htscmdline.c htscmdline.h \
htsurlport.c htsurlport.h
proxytrack_SOURCES = proxy/main.c \
@@ -66,7 +66,7 @@ libhttrack_la_SOURCES = htscore.c htsparse.c htsback.c htscache.c \
htscmdline.c htshelp.c htslib.c htsurlport.c htscoremain.c \
htsname.c htsrobots.c htstools.c htswizard.c \
htsalias.c htsthread.c htsindex.c htsbauth.c \
htsmd5.c htscodec.c htswarc.c htsproxy.c htszlib.c htswrap.c htsconcat.c \
htsmd5.c htscodec.c htswarc.c htssitemap.c htsproxy.c htszlib.c htswrap.c htsconcat.c \
htsmodules.c htscharset.c punycode.c htsencoding.c htssniff.c \
md5.c \
minizip/ioapi.c minizip/mztools.c minizip/unzip.c minizip/zip.c \
@@ -77,7 +77,7 @@ libhttrack_la_SOURCES = htscore.c htsparse.c htsback.c htscache.c \
htshelp.h htsindex.h htslib.h htsurlport.h htsmd5.h \
htsmodules.h htsname.h htsnet.h htssniff.h \
htsopt.h htsrobots.h htsthread.h \
htstools.h htswizard.h htswrap.h htscodec.h htswarc.h htsproxy.h htszlib.h \
htstools.h htswizard.h htswrap.h htscodec.h htswarc.h htssitemap.h htsproxy.h htszlib.h \
htsstrings.h htsarrays.h httrack-library.h \
htscharset.h punycode.h htsencoding.h \
htsentities.h htsentities.sh htsbasiccharsets.sh htscodepages.h \
@@ -87,7 +87,7 @@ libhttrack_la_LIBADD = $(THREADS_LIBS) $(ZLIB_LIBS) $(BROTLI_LIBS) $(ZSTD_LIBS)
libhttrack_la_CFLAGS = $(AM_CFLAGS) -DLIBHTTRACK_EXPORTS -DZLIB_CONST
libhttrack_la_LDFLAGS = $(AM_LDFLAGS) -version-info $(VERSION_INFO)
EXTRA_DIST = httrack.h webhttrack \
EXTRA_DIST = httrack.h htsstats.h webhttrack \
version.rc \
libhttrack.rc \
httrack.rc \

View File

@@ -114,6 +114,10 @@ const char *hts_optalias[][4] = {
"strip [host/pattern=]key1,key2,... from URLs"},
{"cookies-file", "-%K", "param1",
"load extra cookies from a Netscape cookies.txt"},
{"sitemap", "-%m", "single",
"seed the crawl from the start host's sitemap (robots.txt, then "
"/sitemap.xml)"},
{"sitemap-url", "-%mu", "param1", "seed the crawl from this sitemap URL"},
{"warc", "-%r", "single", "write an ISO-28500 WARC/1.1 archive of the crawl"},
{"warc-file", "-%rf", "param1", "write a WARC archive to the given base name"},
{"warc-max-size", "-%rs", "param1",

View File

@@ -142,10 +142,13 @@ struct cache_back_zip_entry {
int compressionMethod;
};
/* Clip: strlcpybuff aborts on overflow, which would take the crawl down on a
corrupt cache carrying a field wider than ours. */
#define ZIP_READFIELD_STRING(line, value, refline, refvalue, refvalue_size) \
do { \
if (line[0] != '\0' && strfield2(line, refline)) { \
strlcpybuff(refvalue, value, refvalue_size); \
(refvalue)[0] = '\0'; \
strlncatbuff(refvalue, value, refvalue_size, (refvalue_size) - 1); \
line[0] = '\0'; \
} \
} while (0)

View File

@@ -1195,6 +1195,12 @@ int cache_legacy_refused_selftest(httrackp *opt, const char *dir) {
/* --- read-side corruption injection --------------------------------------- */
/* 100 'A's: a placeholder header line long enough to be overwritten by a
forged, over-long X-StatusMessage of the same byte length. */
#define CORRUPT_LONG_ETAG \
"AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA" \
"AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA"
/* canary read back intact after each corruption; victim gets the byte surgery
*/
#define CORRUPT_ADR "corrupt.example.com"
@@ -1219,6 +1225,24 @@ static void corrupt_build(httrackp *opt) {
selftest_close(&cache);
}
/* Like corrupt_build, but the victim carries a 100-char Etag placeholder. */
static void corrupt_build_longetag(httrackp *opt) {
cache_back cache;
memset(corrupt_body_a, 'a', sizeof(corrupt_body_a) - 1);
memset(corrupt_body_b, 'b', sizeof(corrupt_body_b) - 1);
remove(reconcile_st_path(opt, "hts-cache/new.zip"));
remove(reconcile_st_path(opt, "hts-cache/old.zip"));
selftest_open_for_write(&cache, opt);
store_entry(opt, &cache, CORRUPT_ADR, "/canary.html", "canary.html", 200,
"OK", "text/html", "utf-8", "", "", "", "", corrupt_body_a,
strlen(corrupt_body_a));
store_entry(opt, &cache, CORRUPT_ADR, "/victim.html", "victim.html", 200,
"OK", "text/html", "utf-8", "", CORRUPT_LONG_ETAG, "", "",
corrupt_body_b, strlen(corrupt_body_b));
selftest_close(&cache);
}
/* Like corrupt_build, but the victim carries a 20-char Etag whose header line
is later overwritten with a forged oversized X-Size (same byte length). */
static void corrupt_build_etag(httrackp *opt) {
@@ -1403,6 +1427,48 @@ static int corrupt_expect_disk_header(httrackp *opt, LLint wantsize,
return fail;
}
/* An over-long field from a foreign cache must clip, not abort: the entry
still reads, clipped to capacity, and the canary survives. */
static int corrupt_expect_victim_clipped(httrackp *opt, size_t wantmsg,
size_t wantlastmod, const char *what) {
cache_back cache;
htsblk v, c;
char BIGSTK lv[HTS_URLMAXSIZE * 2];
char BIGSTK lc[HTS_URLMAXSIZE * 2];
int fail = 0;
selftest_open_for_read(&cache, opt);
lv[0] = lc[0] = '\0';
v = cache_readex(opt, &cache, CORRUPT_ADR, "/victim.html", "", lv, NULL, 1);
if (v.statuscode != 200) {
fprintf(stderr, "%s: %s: status %d, expected 200\n", selftest_tag, what,
v.statuscode);
fail++;
}
if (wantmsg != (size_t) -1 && strlen(v.msg) != wantmsg) {
fprintf(stderr, "%s: %s: msg len %u, expected %u\n", selftest_tag, what,
(unsigned) strlen(v.msg), (unsigned) wantmsg);
fail++;
}
if (wantlastmod != (size_t) -1 && strlen(v.lastmodified) != wantlastmod) {
fprintf(stderr, "%s: %s: lastmodified len %u, expected %u\n", selftest_tag,
what, (unsigned) strlen(v.lastmodified), (unsigned) wantlastmod);
fail++;
}
c = cache_readex(opt, &cache, CORRUPT_ADR, "/canary.html", "", lc, NULL, 1);
if (c.statuscode != 200) {
fprintf(stderr, "%s: %s: canary tainted (status %d)\n", selftest_tag, what,
c.statuscode);
fail++;
}
if (v.adr != NULL)
freet(v.adr);
if (c.adr != NULL)
freet(c.adr);
selftest_close(&cache);
return fail;
}
/* One zip corruption case: build, patch, then check victim+canary in-session.
*/
static int corrupt_case_zip(httrackp *opt, const char *pat, const char *rep,
@@ -1439,6 +1505,31 @@ int cache_corruption_selftest(httrackp *opt, const char *dir) {
failures += corrupt_expect_victim(opt, "Cache Read Error : Read Data",
"garbled deflate stream");
/* A corrupt cache can hold a field wider than ours. Clipping keeps the
entry; aborting would take the crawl down. Overwrite the placeholder Etag
line in place, same byte length, so the zip offsets stay intact. */
corrupt_build_longetag(opt);
corrupt_patch(opt, "Etag: " CORRUPT_LONG_ETAG, 106,
"X-StatusMessage: " /* 17 + 89 = 106 */
"AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA"
"AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA",
1, 1);
failures +=
corrupt_expect_victim_clipped(opt, sizeof(((htsblk *) 0)->msg) - 1,
(size_t) -1, "over-long X-StatusMessage");
/* lastmodified[64] is narrower than msg[80]: one hardcoded clip length
cannot satisfy both. */
corrupt_build_longetag(opt);
corrupt_patch(opt, "Etag: " CORRUPT_LONG_ETAG, 106,
"Last-Modified: " /* 15 + 91 = 106 */
"AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA"
"AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA",
1, 1);
failures += corrupt_expect_victim_clipped(
opt, (size_t) -1, sizeof(((htsblk *) 0)->lastmodified) - 1,
"over-long Last-Modified");
/* An X-Size above INT_MAX is positive as int64 (slips a bare sign check) but
truncates negative in the (int) cast the malloc uses: a wraparound alloc.
cache_add asserts size fits an int, so such a value only reaches the reader

View File

@@ -39,6 +39,7 @@ Please visit our Website: http://www.httrack.com
/* File defs */
#include "htscore.h"
#include "htssitemap.h"
#include "htswarc.h"
/* specific definitions */
@@ -943,6 +944,22 @@ int httpmirror(char *url1, httrackp * opt) {
heap_top()->premier = heap_top_index(); // premier lien, objet-père=objet
heap_top()->precedent = heap_top_index(); // lien précédent
/* --sitemap: queue the sitemap probe just after the seeds, so its URLs are
injected before the crawl gets far. */
hts_sitemap_free(opt); /* an earlier mirror may have left a doc list */
if (opt->sitemap || StringNotEmpty(opt->sitemap_url)) {
char BIGSTK first[HTS_URLMAXSIZE * 2];
const char *const eol = strchr(primary, '\n');
const size_t len = eol != NULL ? (size_t) (eol - primary) : 0;
first[0] = '\0';
if (len > 0 && len < sizeof(first)) {
memcpy(first, primary, len);
first[len] = '\0';
}
hts_sitemap_seed(opt, first);
}
// Initialiser cache
{
opt->state._hts_in_html_parsing = 4;
@@ -1587,11 +1604,21 @@ int httpmirror(char *url1, httrackp * opt) {
stre.maketrack_fp = maketrack_fp;
/* Parse */
if (hts_mirror_check_moved(&str, &stre) != 0) {
XH_uninit;
return -1;
}
{
const int nlinks = opt->lien_tot;
if (hts_mirror_check_moved(&str, &stre) != 0) {
XH_uninit;
return -1;
}
/* A redirect re-queues the target as a fresh link; without carrying
the marking over, a moved sitemap is fetched and then ignored. */
if (opt->sitemap_state != NULL && opt->lien_tot > nlinks &&
hts_sitemap_pending(opt, urladr(), urlfil())) {
hts_sitemap_redirect(opt, urladr(), urlfil(), heap_top()->adr,
heap_top()->fil);
}
}
}
} // if !error
@@ -1609,6 +1636,29 @@ int httpmirror(char *url1, httrackp * opt) {
/* Load file and decode if necessary, after redirect check. */
LOAD_IN_MEMORY_IF_NECESSARY();
/* Sitemap document: turn its <loc> URLs into top-level seeds. They go
through htsAddLink, so the wizard's filters and scope rules decide, and
this link's max depth leaves them the full budget. */
if (opt->sitemap_state != NULL &&
hts_sitemap_pending(opt, urladr(), urlfil())) {
htsmoduleStruct BIGSTK smstr;
int smptr = ptr;
memset(&smstr, 0, sizeof(smstr));
smstr.opt = opt;
smstr.sback = sback;
smstr.cache = &cache;
smstr.hashptr = hashptr;
smstr.numero_passe = numero_passe;
smstr.ptr_ = &smptr; /* scratch: the ingester retargets the wizard */
smstr.addLink = htsAddLink;
smstr.url_host = urladr();
smstr.url_file = urlfil();
smstr.mime = r.contenttype;
hts_sitemap_ingest(opt, &smstr, urladr(), urlfil(), r.adr,
r.adr != NULL && r.size > 0 ? (size_t) r.size : 0);
}
// ------------------------------------------------------
// ok, fichier chargé localement
// ------------------------------------------------------
@@ -1819,6 +1869,9 @@ int httpmirror(char *url1, httrackp * opt) {
if (strnotempty(savename()) == 0) { // pas de chemin de sauvegarde
if (strcmp(urlfil(), "/robots.txt") == 0) { // robots.txt
char BIGSTK sitemaps[8192];
sitemaps[0] = '\0';
if (r.adr) {
char BIGSTK infobuff[8192];
#ifdef IGNORE_RESTRICTIVE_ROBOTS
@@ -1830,7 +1883,8 @@ int httpmirror(char *url1, httrackp * opt) {
#endif
robots_parse(&robots, urladr(), r.adr, r.size, infobuff,
sizeof(infobuff), keep_root);
sizeof(infobuff), keep_root, sitemaps,
sizeof(sitemaps));
if (strnotempty(infobuff)) {
hts_log_print(opt, LOG_INFO,
"Note: robots.txt forbidden links for %s are: %s",
@@ -1840,6 +1894,10 @@ int httpmirror(char *url1, httrackp * opt) {
urladr(), infobuff);
}
}
/* After robots_parse, so the rules this very body carries already
gate the sitemap fetch. Runs even on a failed probe, which is
what falls back to the well-known location. */
hts_sitemap_robots(opt, urladr(), sitemaps);
}
} else if (r.is_write) { // déja sauvé sur disque
/*
@@ -2256,6 +2314,7 @@ int httpmirror(char *url1, httrackp * opt) {
// ending
usercommand(opt, 0, NULL, NULL, NULL, NULL);
warc_close_opt(opt);
hts_sitemap_free(opt);
// désallocation mémoire & buffers
XH_uninit;
@@ -3643,6 +3702,11 @@ HTSEXT_API int copy_htsopt(const httrackp * from, httrackp * to) {
to->warc_cdx = from->warc_cdx;
to->warc_wacz = from->warc_wacz;
if (from->sitemap)
to->sitemap = from->sitemap;
if (StringNotEmpty(from->sitemap_url))
StringCopyS(to->sitemap_url, from->sitemap_url);
if (from->pause_max_ms > 0) {
to->pause_min_ms = from->pause_min_ms;
to->pause_max_ms = from->pause_max_ms;

View File

@@ -1795,6 +1795,26 @@ static int hts_main_internal(int argc, char **argv, httrackp * opt) {
StringCopy(opt->warc_file, WARC_AUTONAME);
}
break;
case 'm': // sitemap / sitemap-url: seed the crawl from sitemaps
if (*(com + 1) == 'u') { // --sitemap-url URL: explicit sitemap
com++;
if ((na + 1 >= argc) || (argv[na + 1][0] == '-')) {
HTS_PANIC_PRINTF(
"Option sitemap-url needs a blank space and a URL");
htsmain_free();
return -1;
}
na++;
if (strlen(argv[na]) >= HTS_URLMAXSIZE) {
HTS_PANIC_PRINTF("Sitemap URL too long");
htsmain_free();
return -1;
}
StringCopy(opt->sitemap_url, argv[na]);
} else { // --sitemap: robots.txt probe, then /sitemap.xml
opt->sitemap = HTS_TRUE;
}
break;
case 'Y': // why: explain the filter verdict for a URL, no crawl
if ((na + 1 >= argc) || (argv[na + 1][0] == '-')) {
HTS_PANIC_PRINTF("Option why needs a blank space and a URL");

View File

@@ -433,7 +433,8 @@ void help_catchurl(const char *dest_path) {
}
// former URL!
{
char BIGSTK finalurl[HTS_URLMAXSIZE * 2];
/* url and dest are each HTS_URLMAXSIZE*2, plus the POSTTOK marker */
char BIGSTK finalurl[HTS_URLMAXSIZE * 4 + 32];
inplace_escape_check_url(dest, sizeof(dest));
snprintf(finalurl, sizeof(finalurl), "%s" POSTTOK "file:%s", url, dest);
@@ -525,6 +526,11 @@ void help(const char *app, int more) {
(" %L <file> add all URL located in this text file (one URL per line)");
infomsg
(" %S <file> add all scan rules located in this text file (one scan rule per line)");
infomsg(" %m seed the crawl from the site's sitemap (robots.txt Sitemap:, "
"then /sitemap.xml); --sitemap-url URL names one explicitly. A "
"sitemap you name, or one the site declares, is fetched even under "
"robots.txt Disallow; only the guessed /sitemap.xml obeys it. The "
"URLs found still pass every filter and scope rule");
infomsg("");
infomsg("Build options:");
infomsg(" NN structure type (0 *original structure, 1+: see below)");

View File

@@ -36,6 +36,7 @@ Please visit our Website: http://www.httrack.com
// Fichier librairie .c
#include "htscore.h"
#include "htssitemap.h"
#include "htswarc.h"
/* specific definitions */
@@ -6019,6 +6020,7 @@ HTSEXT_API httrackp *hts_create_opt(void) {
StringCopy(opt->strip_query, "");
StringCopy(opt->cookies_file, "");
StringCopy(opt->warc_file, "");
StringCopy(opt->sitemap_url, "");
opt->warc_max_size = 0; /* no rotation unless --warc-max-size sets it */
StringCopy(opt->why_url, "");
opt->pause_min_ms = 0;
@@ -6172,6 +6174,8 @@ HTSEXT_API void hts_free_opt(httrackp * opt) {
StringFree(opt->cookies_file);
StringFree(opt->why_url);
StringFree(opt->warc_file);
StringFree(opt->sitemap_url);
hts_sitemap_free(opt); /* backstop: httpmirror's early-return paths */
StringFree(opt->path_html);
StringFree(opt->path_html_utf8);

View File

@@ -547,6 +547,13 @@ struct httrackp {
archive. Tail: ABI */
hts_boolean warc_wacz; /**< --wacz: package archive+index+pages as a WACZ zip
(implies --warc + --warc-cdx). Tail: ABI */
hts_boolean sitemap; /**< --sitemap: probe the start host's robots.txt for
Sitemap: lines, else /sitemap.xml. Tail: ABI */
String sitemap_url; /**< --sitemap-url: sitemap to ingest. Tail: ABI */
/* Live state, not an option: copy_htsopt must leave it alone. It sits here
rather than in htsoptstate because that struct is embedded by value, so
growing it would shift every httrackp field declared after it. */
void *sitemap_state; /**< hts_sitemap_state*, or NULL. Tail: ABI */
};
/* Running statistics for a mirror. */

View File

@@ -147,7 +147,8 @@ static void robots_blob_add(char *blob, size_t blobsize, char marker,
void robots_parse(robots_wizard *robots, const char *adr, const char *body,
size_t bodysize, char *info, size_t infosize,
hts_boolean keep_root_disallow) {
hts_boolean keep_root_disallow, char *sitemaps,
size_t sitemapsize) {
size_t bptr = 0;
int record = 0;
char BIGSTK line[1024];
@@ -156,6 +157,8 @@ void robots_parse(robots_wizard *robots, const char *adr, const char *body,
blob[0] = '\0';
if (info != NULL && infosize > 0)
info[0] = '\0';
if (sitemaps != NULL && sitemapsize > 0)
sitemaps[0] = '\0';
#if DEBUG_ROBOTS
printf("robots.txt dump:\n%s\n", body);
#endif
@@ -172,7 +175,19 @@ void robots_parse(robots_wizard *robots, const char *adr, const char *body,
line[llen - 1] = '\0';
llen--;
}
if (strfield(line, "user-agent:")) {
if (sitemaps != NULL && strfield(line, "sitemap:")) {
// group-independent record (RFC 9309): collected whatever the group
char *a = line + 8;
while (is_realspace(*a))
a++;
/* A line at the buffer limit was truncated: a half URL is not one. */
if (strnotempty(a) && strlen(line) < sizeof(line) - 3 &&
strlen(a) + 2 < sitemapsize - strlen(sitemaps)) {
strlcatbuff(sitemaps, a, sitemapsize);
strlcatbuff(sitemaps, "\n", sitemapsize);
}
} else if (strfield(line, "user-agent:")) {
char *a = line + 11;
while (is_realspace(*a))

View File

@@ -56,10 +56,12 @@ int checkrobots(robots_wizard * robots, const char *adr, const char *fil);
void checkrobots_free(robots_wizard * robots);
int checkrobots_set(robots_wizard * robots, const char *adr, const char *data);
/* Parse robots.txt `body` for `adr`, storing the HTTrack group's rules; `info`
gets a disallow summary, `keep_root_disallow` FALSE drops "Disallow: /". */
gets a disallow summary, `keep_root_disallow` FALSE drops "Disallow: /", and
`sitemaps` (optional) collects the Sitemap: URLs, one per line. */
void robots_parse(robots_wizard *robots, const char *adr, const char *body,
size_t bodysize, char *info, size_t infosize,
hts_boolean keep_root_disallow);
hts_boolean keep_root_disallow, char *sitemaps,
size_t sitemapsize);
#endif
#endif

View File

@@ -59,6 +59,7 @@ Please visit our Website: http://www.httrack.com
#include "htssniff.h"
#include "htscodec.h"
#include "htsproxy.h"
#include "htssitemap.h"
#include "htswarc.h"
#if HTS_USEZLIB
#include "htszlib.h"
@@ -457,6 +458,12 @@ static void basic_selftests(void) {
// link one level up -> a "../" prefix
assertf(lienrelatif(s, sizeof(s), "a.html", "dir/index.html") == 0);
assertf(strcmp(s, "../a.html") == 0);
// an empty current path: the trim used to walk off the front of it, which
// "?x" reaches too because the query pre-pass hands on the part before it
assertf(lienrelatif(s, sizeof(s), "dir/page.html", "") == 0);
assertf(strcmp(s, "dir/page.html") == 0);
assertf(lienrelatif(s, sizeof(s), "dir/page.html", "?x") == 0);
assertf(strcmp(s, "dir/page.html") == 0);
}
}
@@ -1519,7 +1526,7 @@ static int st_hashtable(httrackp *opt, int argc, char **argv) {
size_t i;
for (i = bench[loop].offset; i < (size_t) count;
i += bench[loop].modulus) {
int result;
int result = 0; /* no final else: an unknown type reports failure */
FMT();
if (bench[loop].type == DO_ADD || bench[loop].type == DO_DRY_ADD) {
size_t k;
@@ -1674,6 +1681,22 @@ static int st_copyopt(httrackp *opt, int argc, char **argv) {
if (strcmp(StringBuff(to->warc_file), "run.warc.gz") != 0)
err = 1;
/* sitemap pair: the flag latches on, the URL takes the String deep copy */
from->sitemap = HTS_TRUE;
StringCopy(from->sitemap_url, "http://h.test/sitemap.xml");
to->sitemap = HTS_FALSE;
StringCopy(to->sitemap_url, "");
copy_htsopt(from, to);
if (!to->sitemap ||
strcmp(StringBuff(to->sitemap_url), "http://h.test/sitemap.xml") != 0)
err = 1;
from->sitemap = HTS_FALSE;
StringCopy(from->sitemap_url, "");
copy_htsopt(from, to);
if (!to->sitemap ||
strcmp(StringBuff(to->sitemap_url), "http://h.test/sitemap.xml") != 0)
err = 1;
/* #185 pause pair: copied when enabled (max>0), the 0 sentinel skips */
from->pause_min_ms = 5000;
from->pause_max_ms = 10000;
@@ -3598,7 +3621,7 @@ static int rb_decide(robots_wizard *r, const char *txt, const char *path) {
char host[64];
snprintf(host, sizeof(host), "h%d.example", n++);
robots_parse(r, host, txt, strlen(txt), NULL, 0, HTS_TRUE);
robots_parse(r, host, txt, strlen(txt), NULL, 0, HTS_TRUE, NULL, 0);
return checkrobots(r, host, path);
}
@@ -3675,6 +3698,267 @@ static int st_robots(httrackp *opt, int argc, char **argv) {
return 0;
}
/* Collect the URLs a sitemap scan hands out. */
typedef struct sm_collect {
int n;
char url[8][HTS_URLMAXSIZE];
} sm_collect;
static hts_boolean sm_take(void *arg, const char *url) {
sm_collect *const c = (sm_collect *) arg;
if (c->n < (int) (sizeof(c->url) / sizeof(c->url[0])))
strcpybuff(c->url[c->n], url);
c->n++;
return HTS_TRUE;
}
/* Scan `doc` off a heap buffer with no NUL terminator, so a read past the
declared size is an ASan error rather than a silent pass. */
static int sm_scan(const char *doc, int maxurls, hts_boolean *is_index,
sm_collect *out) {
const size_t len = strlen(doc);
char *raw = malloct(len);
int n;
memset(out, 0, sizeof(*out));
assertf(raw != NULL);
memcpy(raw, doc, len);
n = hts_sitemap_scan(raw, len, maxurls, is_index, sm_take, out);
freet(raw);
return n;
}
static int st_sitemap(httrackp *opt, int argc, char **argv) {
sm_collect c;
hts_boolean idx;
(void) opt;
(void) argc;
(void) argv;
/* A urlset yields its <loc> URLs, in order, unescaped. */
assertf(sm_scan("<?xml version=\"1.0\"?><urlset>"
"<url><loc>http://h.test/a.html</loc></url>"
"<url><loc> https://h.test/b?x=1&amp;y=2\n </loc></url>"
"</urlset>",
100, &idx, &c) == 2);
assertf(!idx);
assertf(strcmp(c.url[0], "http://h.test/a.html") == 0);
assertf(strcmp(c.url[1], "https://h.test/b?x=1&y=2") == 0);
/* A sitemapindex is flagged: its URLs are child sitemaps, not pages. */
assertf(sm_scan("<sitemapindex><sitemap><loc>http://h.test/s2.xml.gz</loc>"
"</sitemap></sitemapindex>",
100, &idx, &c) == 1);
assertf(idx);
/* Root element decides even when the other name appears later as text. */
assertf(sm_scan("<urlset><url><loc>http://h.test/a</loc></url>"
"<!-- sitemapindex --></urlset>",
100, &idx, &c) == 1);
assertf(!idx);
/* Numeric character references, decimal and hex, decode to ASCII. */
assertf(sm_scan("<urlset><loc>http://h.test/a&#63;b&#x3D;c</loc></urlset>",
100, &idx, &c) == 1);
assertf(strcmp(c.url[0], "http://h.test/a?b=c") == 0);
/* A reference decoding to a control byte is dropped: the shared decoder
writes the real character and the URL check refuses it. A reference the
decoder cannot represent (&#0;) stays verbatim, like an unknown entity. */
assertf(sm_scan("<urlset><loc>http://h.test/a&#10;b</loc></urlset>", 100,
&idx, &c) == 0);
assertf(sm_scan("<urlset><loc>http://h.test/a&#9;b</loc></urlset>", 100, &idx,
&c) == 0);
assertf(sm_scan("<urlset><loc>http://h.test/a&#0;b</loc></urlset>", 100, &idx,
&c) == 1);
assertf(strcmp(c.url[0], "http://h.test/a&#0;b") == 0);
/* A comment naming the other root element must not flip the verdict. */
assertf(sm_scan("<!-- <sitemapindex> --><urlset><url>"
"<loc>http://h.test/p</loc></url></urlset>",
100, &idx, &c) == 1);
assertf(!idx);
assertf(sm_scan("<?xml version=\"1.0\"?><!-- <urlset> -->"
"<sitemapindex><loc>http://h.test/s</loc></sitemapindex>",
100, &idx, &c) == 1);
assertf(idx);
/* <location> is not <loc>. */
assertf(sm_scan("<urlset><location>http://h.test/a</location></urlset>", 100,
&idx, &c) == 0);
/* Rejected: relative, non-http scheme, embedded space, empty. */
assertf(sm_scan("<urlset><loc>/a.html</loc><loc>ftp://h.test/a</loc>"
"<loc>javascript:alert(1)</loc>"
"<loc>http://h.test/a b</loc><loc></loc></urlset>",
100, &idx, &c) == 0);
/* The URL length bound: one under fits, exactly at it is dropped rather than
truncated into a different URL. */
{
char BIGSTK doc[HTS_URLMAXSIZE * 2];
char BIGSTK url[HTS_URLMAXSIZE + 1];
size_t i;
strcpybuff(url, "http://h.test/");
for (i = strlen(url); i < HTS_URLMAXSIZE - 1; i++)
url[i] = 'a';
url[i] = '\0';
snprintf(doc, sizeof(doc), "<urlset><loc>%s</loc></urlset>", url);
assertf(sm_scan(doc, 100, &idx, &c) == 1);
url[i] = 'a';
url[i + 1] = '\0';
snprintf(doc, sizeof(doc), "<urlset><loc>%s</loc></urlset>", url);
assertf(sm_scan(doc, 100, &idx, &c) == 0);
}
/* The URL cap stops the scan. */
assertf(sm_scan("<urlset><loc>http://h.test/1</loc><loc>http://h.test/2</loc>"
"<loc>http://h.test/3</loc></urlset>",
2, &idx, &c) == 2);
/* The per-document cap at the value the engine actually uses. */
{
const int many = HTS_SITEMAP_MAX_URLS_DOC + 10;
const size_t cap = (size_t) many * 40 + 32;
char *big = malloct(cap);
size_t off;
int i;
assertf(big != NULL);
off = (size_t) snprintf(big, cap, "<urlset>");
assertf(off < cap);
for (i = 0; i < many; i++) {
const int len =
snprintf(big + off, cap - off, "<loc>http://h.test/%d</loc>", i);
assertf(len > 0 && (size_t) len < cap - off);
off += (size_t) len;
}
memset(&c, 0, sizeof(c));
assertf(hts_sitemap_scan(big, off, HTS_SITEMAP_MAX_URLS_DOC, &idx, sm_take,
&c) == HTS_SITEMAP_MAX_URLS_DOC);
/* The handler count, not just the return: a call site hardcoding a smaller
cap would still return its own argument. */
assertf(c.n == HTS_SITEMAP_MAX_URLS_DOC);
freet(big);
}
#if HTS_USEZLIB
/* A highly compressible document decodes without running away: the ratio
budget cannot bind (deflate tops out near 1032:1), so this pins the
decompression path itself rather than the 64 MiB ceiling. */
{
const char *const one = "<url><loc>http://h.test/bomb</loc></url>";
const size_t reps = 40000;
size_t xlen = 8 + reps * strlen(one) + 10, i;
char *x = malloct(xlen + 1);
uLongf zlen;
char *z;
z_stream zs;
assertf(x != NULL);
{
size_t w = (size_t) snprintf(x, xlen, "<urlset>");
int len;
assertf(w < xlen);
for (i = 0; i < reps; i++) {
len = snprintf(x + w, xlen - w, "%s", one);
assertf(len > 0 && (size_t) len < xlen - w);
w += (size_t) len;
}
len = snprintf(x + w, xlen - w, "</urlset>");
assertf(len > 0 && (size_t) len < xlen - w);
w += (size_t) len;
xlen = w;
}
zlen = compressBound((uLong) xlen) + 32;
z = malloct((size_t) zlen);
assertf(z != NULL);
memset(&zs, 0, sizeof(zs));
assertf(deflateInit2(&zs, 9, Z_DEFLATED, 16 + MAX_WBITS, 8,
Z_DEFAULT_STRATEGY) == Z_OK);
zs.next_in = (const Bytef *) x;
zs.avail_in = (uInt) xlen;
zs.next_out = (Bytef *) z;
zs.avail_out = (uInt) zlen;
assertf(deflate(&zs, Z_FINISH) == Z_STREAM_END);
zlen = (uLongf) zs.total_out;
deflateEnd(&zs);
/* well over the 4096:1 budget's 1 MiB floor, and far under the 64 MiB cap
*/
assertf(xlen > 1024 * 1024 && (size_t) zlen < xlen / 100);
memset(&c, 0, sizeof(c));
assertf(hts_sitemap_scan(z, (size_t) zlen, 10, &idx, sm_take, &c) == 10);
assertf(strcmp(c.url[0], "http://h.test/bomb") == 0);
freet(z);
freet(x);
}
#endif
/* An unterminated <loc> at end of buffer must not read past it. */
assertf(sm_scan("<urlset><loc>http://h.test/a", 100, &idx, &c) == 0);
assertf(sm_scan("<urlset><lo", 100, &idx, &c) == 0);
#if HTS_USEZLIB
/* A gzip-framed document is decompressed before scanning. */
{
const char *const xml =
"<urlset><url><loc>http://h.test/gz.html</loc></url></urlset>";
uLongf zlen = compressBound((uLong) strlen(xml)) + 32;
char *z = malloct((size_t) zlen);
z_stream zs;
assertf(z != NULL);
memset(&zs, 0, sizeof(zs));
assertf(deflateInit2(&zs, 9, Z_DEFLATED, 16 + MAX_WBITS, 8,
Z_DEFAULT_STRATEGY) == Z_OK);
zs.next_in = (const Bytef *) xml;
zs.avail_in = (uInt) strlen(xml);
zs.next_out = (Bytef *) z;
zs.avail_out = (uInt) zlen;
assertf(deflate(&zs, Z_FINISH) == Z_STREAM_END);
zlen = (uLongf) zs.total_out;
deflateEnd(&zs);
memset(&c, 0, sizeof(c));
assertf(hts_sitemap_scan(z, (size_t) zlen, 100, &idx, sm_take, &c) == 1);
assertf(strcmp(c.url[0], "http://h.test/gz.html") == 0);
/* Truncated gzip: refused, not scanned as plain text. */
memset(&c, 0, sizeof(c));
assertf(hts_sitemap_scan(z, 4, 100, &idx, sm_take, &c) == -1);
freet(z);
}
#endif
/* robots.txt: only Sitemap: records, comments stripped, case-insensitive,
and group-independent (no User-agent line needed). */
/* robots_parse collects Sitemap: whatever the user-agent group, strips the
comment and keeps the rules working alongside it. */
{
const char *const txt = "User-agent: *\nDisallow: /x\n"
"SITEMAP: http://h.test/s1.xml # first\n"
"Sitemapper: http://h.test/no.xml\n"
"Sitemap:\thttps://h.test/s2.xml\n";
char BIGSTK maps[1024];
robots_wizard rb;
memset(&rb, 0, sizeof(rb));
robots_parse(&rb, "h.test", txt, strlen(txt), NULL, 0, HTS_TRUE, maps,
sizeof(maps));
assertf(strcmp(maps, "http://h.test/s1.xml\nhttps://h.test/s2.xml\n") == 0);
assertf(checkrobots(&rb, "h.test", "/x") == -1);
checkrobots_free(&rb);
}
printf("sitemap self-test OK\n");
return 0;
}
/* Connected stream pair over loopback; Windows has no socketpair(). */
static int st_socketpair(T_SOC sv[2]) {
struct sockaddr_in sa;
@@ -3772,21 +4056,6 @@ static int st_ftpuser(httrackp *opt, int argc, char **argv) {
return 0;
}
/* Bounded substring search (records carry NUL bytes; strstr won't do). */
static const char *warc_memstr(const char *hay, const char *needle,
size_t haylen, size_t nlen) {
if (nlen == 0 || haylen < nlen)
return NULL;
{
size_t i;
for (i = 0; i + nlen <= haylen; i++) {
if (memcmp(hay + i, needle, nlen) == 0)
return hay + i;
}
}
return NULL;
}
/* Slurp a whole file into a malloc'd buffer; sets *len. NULL on error. */
static unsigned char *warc_slurp(const char *path, size_t *len) {
FILE *f = FOPEN(path, "rb");
@@ -3864,6 +4133,13 @@ static unsigned char *warc_next_member(const unsigned char **in,
Content-Length == block length, the \r\n\r\n trailer intact, the response
body round-trips, and the hop-by-hop Transfer-Encoding is dropped (a real
Content-Encoding is kept verbatim; see warc-verbatim). */
/* Argument order kept for the existing call sites; the search itself is the
shared hts_memstr. */
static const char *warc_memstr(const char *hay, const char *needle,
size_t haylen, size_t nlen) {
return hts_memstr(hay, haylen, needle, nlen);
}
static int st_warc(httrackp *opt, int argc, char **argv) {
char path[HTS_URLMAXSIZE];
warc_writer *w;
@@ -5283,6 +5559,8 @@ static const struct selftest_entry {
st_contentcodings},
{"robots", "", "robots.txt RFC 9309 Allow/Disallow precedence self-test",
st_robots},
{"sitemap", "",
"sitemap <loc> extraction, caps and robots.txt Sitemap:", st_sitemap},
{"ftp-line", "", "get_ftp_line bounds a hostile FTP reply line",
st_ftpline},
{"ftp-userpass", "", "ftp_split_userpass bounds URL userinfo", st_ftpuser},

615
src/htssitemap.c Normal file
View File

@@ -0,0 +1,615 @@
/* ------------------------------------------------------------ */
/*
HTTrack Website Copier, Offline Browser for Windows and Unix
Copyright (C) 2026 Xavier Roche and other contributors
SPDX-License-Identifier: GPL-3.0-or-later
This program is free software: you can redistribute it and/or modify
it under the terms of the GNU General Public License as published by
the Free Software Foundation, either version 3 of the License, or
(at your option) any later version.
This program is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
GNU General Public License for more details.
You should have received a copy of the GNU General Public License
along with this program. If not, see <http://www.gnu.org/licenses/>.
Ethical use: we kindly ask that you NOT use this software to harvest email
addresses or to collect any other private information about people. Doing so
would dishonor our work and waste the many hours we have spent on it.
Please visit our Website: http://www.httrack.com
*/
/* ------------------------------------------------------------ */
/* File: sitemap ingestion (sitemaps.org 0.9) */
/* Author: Xavier Roche */
/* ------------------------------------------------------------ */
#define HTS_INTERNAL_BYTECODE
#include "htscore.h"
#include "htssitemap.h"
#include "htsbase.h"
#include "htscodec.h"
#include "htsencoding.h"
#include "htsfilters.h"
#include "htshash.h"
#include "htsmodules.h"
#include "htslib.h"
#include "htsrobots.h"
#include "htssafe.h"
#include "htstools.h"
#include <ctype.h>
#include <string.h>
/* One queued sitemap document awaiting ingestion. */
typedef struct sitemap_doc {
char adr[HTS_URLMAXSIZE];
char fil[HTS_URLMAXSIZE];
int level;
hts_sitemap_source src;
hts_boolean done;
struct sitemap_doc *next;
} sitemap_doc;
struct hts_sitemap_state {
sitemap_doc *docs;
int ndocs; /* documents queued, capped by HTS_SITEMAP_MAX_DOCS */
int nurls; /* URLs seeded, capped by HTS_SITEMAP_MAX_URLS_TOTAL */
hts_boolean probe_done; /* the robots.txt probe has been answered */
hts_boolean fallback_done; /* the /sitemap.xml fallback was already queued */
/* The crawl's own start URL. Seeded URLs are judged against it, so a site
cannot widen a subtree crawl by putting its sitemap at the root. */
char anchor_adr[HTS_URLMAXSIZE];
char anchor_fil[HTS_URLMAXSIZE];
};
typedef struct hts_sitemap_state hts_sitemap_state;
/* --------------------------------------------------------------------- */
/* Document parsing (no engine state: fuzzable and self-testable) */
/* --------------------------------------------------------------------- */
/* Accept only an absolute http(s) URL with no space or control byte. */
static hts_boolean sitemap_url_ok(const char *url) {
const char *p;
if (!strfield(url, "http://") && !strfield(url, "https://"))
return HTS_FALSE;
for (p = url; *p != '\0'; p++) {
if ((unsigned char) *p <= ' ' || (unsigned char) *p == 0x7f)
return HTS_FALSE;
}
return HTS_TRUE;
}
/* Skip to the character after the next '>' at or after p, or NULL. */
static const char *sitemap_tag_end(const char *p, const char *end) {
while (p < end && *p != '>')
p++;
return p < end ? p + 1 : NULL;
}
/* HTS_TRUE when the document's root element is `name`. Skips the XML
declaration, comments and processing instructions first, so a comment
mentioning the other root element cannot decide the document type. */
static hts_boolean sitemap_root_is(const char *doc, size_t size,
const char *name) {
const size_t nlen = strlen(name);
size_t i = 0;
if (size >= 3 && memcmp(doc, "\xef\xbb\xbf", 3) == 0)
i = 3; /* UTF-8 BOM */
while (i < size) {
if (isspace((unsigned char) doc[i])) {
i++;
} else if (doc[i] != '<') {
return HTS_FALSE; /* character data before any element: not XML */
} else if (i + 4 <= size && memcmp(doc + i, "<!--", 4) == 0) {
const char *const e = hts_memstr(doc + i, size - i, "-->", 3);
if (e == NULL)
return HTS_FALSE;
i = (size_t) (e - doc) + 3;
} else if (i + 2 <= size && (doc[i + 1] == '?' || doc[i + 1] == '!')) {
while (i < size && doc[i] != '>')
i++;
i++;
} else {
size_t j = i + 1;
/* an optional namespace prefix: <sm:sitemapindex> is the same element */
while (j < size && doc[j] != ':' && doc[j] != '>' &&
!isspace((unsigned char) doc[j]))
j++;
if (j >= size || doc[j] != ':')
j = i + 1;
else
j++;
return j + nlen <= size && memcmp(doc + j, name, nlen) == 0 &&
(j + nlen == size ||
isspace((unsigned char) doc[j + nlen]) ||
doc[j + nlen] == '>' || doc[j + nlen] == '/')
? HTS_TRUE
: HTS_FALSE;
}
}
return HTS_FALSE;
}
/* Decompress a gzip-framed body into a fresh buffer. The 64 MiB cap is what
binds in practice; deflate tops out near 1032:1, so the tree's codec budget
only matters as the shared policy for a coding that could go further. */
static char *sitemap_gunzip(const char *body, size_t size, size_t *outsize) {
const LLint budget = hts_codec_maxout((LLint) size);
size_t cap = budget < (LLint) HTS_SITEMAP_MAX_BYTES
? (size_t) budget
: (size_t) HTS_SITEMAP_MAX_BYTES;
char *out;
size_t n;
if (cap == 0)
return NULL;
out = malloct(cap + 1);
if (out == NULL)
return NULL;
n = hts_codec_head(HTS_CODEC_DEFLATE, body, size, out, cap);
if (n == 0) {
freet(out);
return NULL;
}
out[n] = '\0';
*outsize = n;
return out;
}
int hts_sitemap_scan(const char *body, size_t size, int maxurls,
hts_boolean *is_index, hts_sitemap_handler handler,
void *arg) {
char *unpacked = NULL;
const char *doc;
const char *end;
const char *p;
int n = 0;
if (is_index != NULL)
*is_index = HTS_FALSE;
if (body == NULL || size < 2 || handler == NULL)
return 0;
/* Content-Encoding gzip is undone upstream; only the container is left. */
if ((unsigned char) body[0] == 0x1f && (unsigned char) body[1] == 0x8b) {
unpacked = sitemap_gunzip(body, size, &size);
if (unpacked == NULL)
return -1;
doc = unpacked;
} else {
if (size > (size_t) HTS_SITEMAP_MAX_BYTES)
size = (size_t) HTS_SITEMAP_MAX_BYTES;
doc = body;
}
end = doc + size;
/* Set before the first callback: the handler reads the verdict. */
if (is_index != NULL)
*is_index = sitemap_root_is(doc, size, "sitemapindex");
for (p = doc; n < maxurls;) {
const char *loc = hts_memstr(p, (size_t) (end - p), "<loc", 4);
const char *val;
const char *stop;
size_t len;
char BIGSTK url[HTS_URLMAXSIZE];
if (loc == NULL)
break;
/* "<loc>" or "<loc xmlns:..>", never "<location>" */
if (loc + 4 >= end || (loc[4] != '>' && !isspace((unsigned char) loc[4]))) {
p = loc + 4;
continue;
}
val = sitemap_tag_end(loc + 4, end);
if (val == NULL)
break;
for (stop = val; stop < end && *stop != '<'; stop++)
;
/* No closing tag: truncated document, so the value may be a partial URL. */
if (stop == end)
break;
p = stop;
while (val < stop && isspace((unsigned char) *val))
val++;
while (stop > val && isspace((unsigned char) *(stop - 1)))
stop--;
len = (size_t) (stop - val);
/* Overflow-safe: the untrusted length alone against the room left. */
if (len == 0 || len >= sizeof(url))
continue;
memcpy(url, val, len);
url[len] = '\0';
/* hts_unescapeEntities decodes in place and tolerates src == dest; a
reference to a control byte survives as one and sitemap_url_ok drops it.
*/
if (hts_unescapeEntities(url, url, sizeof(url)) != 0 ||
!sitemap_url_ok(url))
continue;
n++;
if (!handler(arg, url))
break;
}
if (unpacked != NULL)
freet(unpacked);
return n;
}
/* --------------------------------------------------------------------- */
/* Engine glue */
/* --------------------------------------------------------------------- */
static hts_sitemap_state *sitemap_get_state(httrackp *opt) {
if (opt->sitemap_state == NULL)
opt->sitemap_state = calloct(1, sizeof(hts_sitemap_state));
return (hts_sitemap_state *) opt->sitemap_state;
}
static sitemap_doc *sitemap_find(httrackp *opt, const char *adr,
const char *fil) {
hts_sitemap_state *const st = (hts_sitemap_state *) opt->sitemap_state;
sitemap_doc *d;
if (st == NULL)
return NULL;
for (d = st->docs; d != NULL; d = d->next) {
if (strfield2(d->adr, adr) && strcmp(d->fil, fil) == 0)
return d;
}
return NULL;
}
/* Who asked for this document decides how far it is gated. The wizard proper
is not usable here: it wants a referring link, and its up/down travel rules
would judge a child sitemap against the parent sitemap's directory. */
static hts_boolean sitemap_fetch_allowed(httrackp *opt, const char *adr,
const char *fil,
hts_sitemap_source src) {
/* adr and fil are each capped just under HTS_URLMAXSIZE, and lfull prefixes
a scheme and a slash on top of both: 2 * HTS_URLMAXSIZE does not fit. */
char BIGSTK l[HTS_URLMAXSIZE * 2 + 16], lfull[HTS_URLMAXSIZE * 2 + 16];
int jokdepth = 0, jok;
hts_boolean refused;
/* The user naming a sitemap is the same intent as naming a start URL, which
the wizard admits unconditionally. */
if (src == HTS_SITEMAP_SRC_USER)
return HTS_TRUE;
strcpybuff(l, jump_identification_const(adr));
if (*fil != '/')
strcatbuff(l, "/");
strcatbuff(l, fil);
strcpybuff(lfull, link_has_authority(adr) ? "" : "http://");
strcatbuff(lfull, adr);
if (*fil != '/')
strcatbuff(lfull, "/");
strcatbuff(lfull, fil);
jok = fa_strjoker_dual(0, *opt->filters.filters, *opt->filters.filptr, lfull,
l, NULL, NULL, &jokdepth);
refused = (jok == -1) ? HTS_TRUE : HTS_FALSE;
if (refused) {
hts_log_print(opt, LOG_NOTICE, "Sitemap: filter rule #%d refuses %s%s",
jokdepth + 1, adr, fil);
return HTS_FALSE;
}
/* A Sitemap: line, or a sitemapindex entry, is the site inviting the fetch;
a Disallow elsewhere in the same file does not retract it. The well-known
location is only ever a guess, so there a Disallow wins. */
if (src == HTS_SITEMAP_SRC_GUESSED &&
hts_robots_forbids(opt, adr, fil, (jok != 0) ? HTS_TRUE : HTS_FALSE,
refused)) {
hts_log_print(opt, LOG_NOTICE, "Sitemap: robots.txt forbids %s%s", adr,
fil);
return HTS_FALSE;
}
return HTS_TRUE;
}
/* Record the link with save="" so the body stays in memory: a sitemap is
ingested, never mirrored. */
static hts_boolean sitemap_queue_(httrackp *opt, const char *adr,
const char *fil, int level,
hts_sitemap_source src, hts_boolean link_it) {
hts_sitemap_state *const st = sitemap_get_state(opt);
sitemap_doc *d;
if (st == NULL)
return HTS_FALSE;
if (st->ndocs >= HTS_SITEMAP_MAX_DOCS || level > HTS_SITEMAP_MAX_LEVEL) {
hts_log_print(opt, LOG_WARNING, "Sitemap: cap reached, skipping %s%s", adr,
fil);
return HTS_FALSE;
}
if (strlen(adr) >= sizeof(d->adr) || strlen(fil) >= sizeof(d->fil))
return HTS_FALSE;
if (sitemap_find(opt, adr, fil) != NULL)
return HTS_FALSE;
if (!sitemap_fetch_allowed(opt, adr, fil, src))
return HTS_FALSE;
d = calloct(1, sizeof(sitemap_doc));
if (d == NULL)
return HTS_FALSE;
strcpybuff(d->adr, adr);
strcpybuff(d->fil, fil);
d->level = level;
d->src = src;
d->next = st->docs;
st->docs = d;
st->ndocs++;
if (!link_it)
return HTS_TRUE;
if (!hts_record_link(opt, adr, fil, "", "", "", NULL))
return HTS_FALSE;
heap_top()->testmode = 0;
heap_top()->link_import = 0;
heap_top()->depth = opt->depth + 1;
heap_top()->pass2 = 0;
heap_top()->retry = opt->retry;
heap_top()->premier = heap_top_index();
heap_top()->precedent = heap_top_index();
hts_log_print(opt, LOG_INFO, "Sitemap: queued %s%s", adr, fil);
return HTS_TRUE;
}
static hts_boolean sitemap_queue(httrackp *opt, const char *adr,
const char *fil, int level,
hts_sitemap_source src) {
return sitemap_queue_(opt, adr, fil, level, src, HTS_TRUE);
}
void hts_sitemap_redirect(httrackp *opt, const char *adr, const char *fil,
const char *newadr, const char *newfil) {
sitemap_doc *const d = sitemap_find(opt, adr, fil);
if (d == NULL || d->done)
return;
d->done = HTS_TRUE; /* the body lives at the target now */
/* The engine already queued the target link, so only the marking moves. */
(void) sitemap_queue_(opt, newadr, newfil, d->level, d->src, HTS_FALSE);
hts_log_print(opt, LOG_NOTICE, "Sitemap: %s%s redirects to %s%s", adr, fil,
newadr, newfil);
}
void hts_sitemap_seed(httrackp *opt, const char *starturl) {
char BIGSTK url[HTS_URLMAXSIZE * 2];
lien_adrfil af;
if (StringNotEmpty(opt->sitemap_url)) {
if (strlen(StringBuff(opt->sitemap_url)) >= sizeof(url)) {
hts_log_print(opt, LOG_ERROR, "Sitemap URL too long");
} else {
strcpybuff(url, StringBuff(opt->sitemap_url));
if (strstr(url, ":/") == NULL)
hts_log_print(opt, LOG_ERROR, "Sitemap URL must be absolute: %s", url);
else if (ident_url_absolute(url, &af) >= 0)
(void) sitemap_queue(opt, af.adr, af.fil, 0, HTS_SITEMAP_SRC_USER);
}
}
if (starturl == NULL || starturl[0] == '\0' ||
strlen(starturl) >= sizeof(url))
return;
strcpybuff(url, starturl);
if (ident_url_absolute(url, &af) < 0)
return;
{
hts_sitemap_state *const st = sitemap_get_state(opt);
if (st != NULL && strlen(af.adr) < sizeof(st->anchor_adr) &&
strlen(af.fil) < sizeof(st->anchor_fil)) {
strcpybuff(st->anchor_adr, af.adr);
strcpybuff(st->anchor_fil, af.fil);
}
}
if (!opt->sitemap)
return;
/* Answered in hts_sitemap_robots, once the parsed rules are installed. */
if (hts_record_link(opt, af.adr, "/robots.txt", "", "", "", NULL)) {
heap_top()->testmode = 0;
heap_top()->link_import = 0;
heap_top()->depth = 0;
heap_top()->pass2 = 0;
heap_top()->retry = opt->retry;
heap_top()->premier = heap_top_index();
heap_top()->precedent = heap_top_index();
/* Claim the host so the parser does not queue robots.txt a second time. */
if (opt->robotsptr != NULL)
(void) checkrobots_set((robots_wizard *) opt->robotsptr, af.adr, "");
}
}
void hts_sitemap_robots(httrackp *opt, const char *adr, const char *sitemaps) {
hts_sitemap_state *const st = (hts_sitemap_state *) opt->sitemap_state;
int queued = 0;
if (st == NULL || !opt->sitemap || st->probe_done ||
!strfield2(st->anchor_adr, adr))
return;
st->probe_done = HTS_TRUE;
if (sitemaps != NULL) {
const char *p = sitemaps;
while (*p != '\0') {
const char *const eol = strchr(p, '\n');
const size_t len = eol != NULL ? (size_t) (eol - p) : strlen(p);
char BIGSTK line[HTS_URLMAXSIZE];
lien_adrfil af;
if (len > 0 && len < sizeof(line)) {
memcpy(line, p, len);
line[len] = '\0';
/* Same host: a Sitemap: line must not aim the fetcher elsewhere. */
if (sitemap_url_ok(line) && ident_url_absolute(line, &af) >= 0 &&
strfield2(af.adr, adr) &&
sitemap_queue(opt, af.adr, af.fil, 0, HTS_SITEMAP_SRC_DECLARED))
queued++;
}
if (eol == NULL)
break;
p = eol + 1;
}
}
if (queued == 0 && !st->fallback_done) {
st->fallback_done = HTS_TRUE;
if (sitemap_queue(opt, adr, "/sitemap.xml", 0, HTS_SITEMAP_SRC_GUESSED))
queued++;
}
hts_log_print(opt, LOG_NOTICE, "Sitemap: %d sitemap(s) queued for %s", queued,
adr);
}
hts_boolean hts_sitemap_pending(httrackp *opt, const char *adr,
const char *fil) {
const sitemap_doc *const d = sitemap_find(opt, adr, fil);
return d != NULL && !d->done ? HTS_TRUE : HTS_FALSE;
}
/* Handler context: seeding URLs from one document. */
typedef struct sitemap_ingest_ctx {
httrackp *opt;
htsmoduleStruct *str;
const char *adr; /* host of the document being ingested */
int level;
hts_boolean is_index;
int accepted; /* URLs seeded or documents queued, not merely parsed */
} sitemap_ingest_ctx;
/* A <loc> of a <urlset>: hand it to the wizard as a top-level seed.
The wizard is pointed at the crawl's own start URL, not at the sitemap: the
site picks where its sitemap lives, so anchoring travel there would let a
root sitemap widen a subtree crawl to the whole host. The URL then becomes
its own anchor, exactly as a command-line seed does. */
static hts_boolean sitemap_seed_url(void *arg, const char *url) {
sitemap_ingest_ctx *const c = (sitemap_ingest_ctx *) arg;
httrackp *const opt = c->opt;
hts_sitemap_state *const st = sitemap_get_state(opt);
char BIGSTK buff[HTS_URLMAXSIZE];
int before;
if (st == NULL || st->nurls >= HTS_SITEMAP_MAX_URLS_TOTAL) {
hts_log_print(opt, LOG_WARNING,
"Sitemap: URL cap reached, ignoring the rest");
return HTS_FALSE;
}
/* strcpybuff aborts rather than truncating: never feed it unchecked input. */
if (strlen(url) >= sizeof(buff))
return HTS_TRUE;
st->nurls++;
strcpybuff(buff, url);
before = opt->lien_tot;
if (htsAddLink(c->str, buff))
c->accepted++;
if (opt->lien_tot > before)
heap_top()->premier = heap_top_index(); /* a seed anchors on itself */
return HTS_TRUE;
}
/* A <loc> of a <sitemapindex>: cross-host children are dropped, so a hostile
sitemap cannot aim the fetcher elsewhere. */
static hts_boolean sitemap_seed_child(void *arg, const char *url) {
sitemap_ingest_ctx *const c = (sitemap_ingest_ctx *) arg;
char BIGSTK buff[HTS_URLMAXSIZE];
lien_adrfil af;
if (strlen(url) >= sizeof(buff))
return HTS_TRUE;
strcpybuff(buff, url);
if (ident_url_absolute(buff, &af) < 0)
return HTS_TRUE;
if (!strfield2(af.adr, c->adr)) {
hts_log_print(c->opt, LOG_WARNING,
"Sitemap: ignoring off-host child sitemap %s%s", af.adr,
af.fil);
return HTS_TRUE;
}
if (sitemap_queue(c->opt, af.adr, af.fil, c->level + 1,
HTS_SITEMAP_SRC_DECLARED))
c->accepted++;
return HTS_TRUE;
}
/* The scan classifies the document before the first callback. */
static hts_boolean sitemap_seed_any(void *arg, const char *url) {
sitemap_ingest_ctx *const c = (sitemap_ingest_ctx *) arg;
return c->is_index ? sitemap_seed_child(arg, url)
: sitemap_seed_url(arg, url);
}
void hts_sitemap_ingest(httrackp *opt, htsmoduleStruct *str, const char *adr,
const char *fil, const char *body, size_t size) {
sitemap_doc *const d = sitemap_find(opt, adr, fil);
hts_sitemap_state *const st = (hts_sitemap_state *) opt->sitemap_state;
sitemap_ingest_ctx ctx;
int n, anchor, saved_depth;
if (d == NULL || d->done)
return;
d->done = HTS_TRUE;
/* str->ptr_ is a scratch int owned by the caller, so nothing else moves. */
anchor = *str->ptr_;
if (st != NULL && st->anchor_adr[0] != '\0' && opt->hash != NULL) {
const int i = hash_read((const hash_struct *) opt->hash, st->anchor_adr,
st->anchor_fil, 1);
if (i >= 0)
anchor = i;
}
*str->ptr_ = anchor;
/* Borrow the anchor's position but keep a seed's full depth budget. */
saved_depth = heap(anchor)->depth;
heap(anchor)->depth = opt->depth + 1;
ctx.opt = opt;
ctx.str = str;
ctx.adr = adr;
ctx.level = d->level;
ctx.is_index = HTS_FALSE;
ctx.accepted = 0;
n = hts_sitemap_scan(body, size, HTS_SITEMAP_MAX_URLS_DOC, &ctx.is_index,
sitemap_seed_any, &ctx);
heap(anchor)->depth = saved_depth;
if (n < 0) {
hts_log_print(opt, LOG_ERROR, "Sitemap: could not decompress %s%s", adr,
fil);
return;
}
if (ctx.is_index)
hts_log_print(opt, LOG_NOTICE,
"Sitemap: %d of %d child sitemap(s) listed by %s%s",
ctx.accepted, n, adr, fil);
else
hts_log_print(opt, LOG_NOTICE, "Sitemap: %d of %d URL(s) added from %s%s",
ctx.accepted, n, adr, fil);
}
void hts_sitemap_free(httrackp *opt) {
hts_sitemap_state *const st = (hts_sitemap_state *) opt->sitemap_state;
if (st == NULL)
return;
while (st->docs != NULL) {
sitemap_doc *const next = st->docs->next;
freet(st->docs);
st->docs = next;
}
freet(opt->sitemap_state);
opt->sitemap_state = NULL;
}

108
src/htssitemap.h Normal file
View File

@@ -0,0 +1,108 @@
/* ------------------------------------------------------------ */
/*
HTTrack Website Copier, Offline Browser for Windows and Unix
Copyright (C) 2026 Xavier Roche and other contributors
SPDX-License-Identifier: GPL-3.0-or-later
This program is free software: you can redistribute it and/or modify
it under the terms of the GNU General Public License as published by
the Free Software Foundation, either version 3 of the License, or
(at your option) any later version.
This program is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
GNU General Public License for more details.
You should have received a copy of the GNU General Public License
along with this program. If not, see <http://www.gnu.org/licenses/>.
Ethical use: we kindly ask that you NOT use this software to harvest email
addresses or to collect any other private information about people. Doing so
would dishonor our work and waste the many hours we have spent on it.
Please visit our Website: http://www.httrack.com
*/
/* ------------------------------------------------------------ */
/* HTTrack sitemap ingestion (sitemaps.org 0.9). Internal, not installed.
Reads <urlset>/<sitemapindex> documents, plain or gzip-framed, and feeds
their <loc> URLs to the crawl as top-level seeds. The whole input is
attacker-controlled, so every entry point below is capped. */
/* ------------------------------------------------------------ */
#ifndef HTS_SITEMAP_DEFH
#define HTS_SITEMAP_DEFH
#include "htsdefines.h"
#include "htsopt.h"
#ifdef __cplusplus
extern "C" {
#endif
/* Caps. sitemaps.org allows 50000 URLs and 50 MB uncompressed per document;
the byte cap sits above that so a conformant sitemap always fits. */
#define HTS_SITEMAP_MAX_URLS_DOC 50000 /* <loc> per document */
#define HTS_SITEMAP_MAX_URLS_TOTAL 200000 /* <loc> per mirror */
#define HTS_SITEMAP_MAX_DOCS 256 /* documents per mirror */
#define HTS_SITEMAP_MAX_LEVEL 4 /* sitemapindex nesting */
#define HTS_SITEMAP_MAX_BYTES (64 * 1024 * 1024) /* decompressed document */
/* Who asked for a sitemap document, which decides how far its fetch is gated.
The user naming one is the same intent as a start URL; a site declaring one
invites the fetch; the well-known location is only ever our guess. */
typedef enum {
HTS_SITEMAP_SRC_USER, /**< --sitemap-url */
HTS_SITEMAP_SRC_DECLARED, /**< a Sitemap: line or a sitemapindex entry */
HTS_SITEMAP_SRC_GUESSED /**< the /sitemap.xml fallback */
} hts_sitemap_source;
/* Per-URL handler; returning HTS_FALSE stops the scan. */
typedef hts_boolean (*hts_sitemap_handler)(void *arg, const char *url);
/* Scan one sitemap document, plain or gzip-framed, handing every acceptable
absolute http(s) <loc> URL to `handler`. Stops after `maxurls` URLs, or when
the handler refuses. `is_index` (optional) reports a <sitemapindex>, whose
URLs are child sitemaps rather than pages. Returns the number of URLs handed
out, or -1 when the document could not be decompressed within the caps. */
int hts_sitemap_scan(const char *body, size_t size, int maxurls,
hts_boolean *is_index, hts_sitemap_handler handler,
void *arg);
/* --- Engine glue (needs a live httrackp). --- */
/* Queue the first sitemap document of the mirror: the explicit --sitemap-url,
or the start host's /robots.txt probe for --sitemap. `starturl` is the first
command-line seed. No-op when neither option is set. */
void hts_sitemap_seed(httrackp *opt, const char *starturl);
/* Act on the start host's robots.txt once its rules are installed: queue the
Sitemap: URLs it names (newline-separated, from robots_parse), or the
well-known /sitemap.xml when it names none. No-op unless --sitemap. */
void hts_sitemap_robots(httrackp *opt, const char *adr, const char *sitemaps);
/* Carry the sitemap marking of (adr,fil) over to the target of a redirect the
engine has already queued, so a moved sitemap is still ingested. */
void hts_sitemap_redirect(httrackp *opt, const char *adr, const char *fil,
const char *newadr, const char *newfil);
/* HTS_TRUE when (adr,fil) is a queued sitemap document awaiting ingestion. */
hts_boolean hts_sitemap_pending(httrackp *opt, const char *adr,
const char *fil);
/* Ingest a fetched sitemap document (or the robots.txt probe): seed its URLs
through the wizard via htsAddLink, and queue nested sitemaps. `str` supplies
the parser context of the document being processed. */
void hts_sitemap_ingest(httrackp *opt, htsmoduleStruct *str, const char *adr,
const char *fil, const char *body, size_t size);
/* Release the ingestion state held in opt (NULL-safe, idempotent). */
void hts_sitemap_free(httrackp *opt);
#ifdef __cplusplus
}
#endif
#endif

63
src/htsstats.h Normal file
View File

@@ -0,0 +1,63 @@
/* ------------------------------------------------------------ */
/*
HTTrack Website Copier, Offline Browser for Windows and Unix
Copyright (C) 1998 Xavier Roche and other contributors
SPDX-License-Identifier: GPL-3.0-or-later
This program is free software: you can redistribute it and/or modify
it under the terms of the GNU General Public License as published by
the Free Software Foundation, either version 3 of the License, or
(at your option) any later version.
This program is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
GNU General Public License for more details.
You should have received a copy of the GNU General Public License
along with this program. If not, see <http://www.gnu.org/licenses/>.
Ethical use: we kindly ask that you NOT use this software to harvest email
addresses or to collect any other private information about people. Doing so
would dishonor our work and waste the many hours we have spent on it.
Please visit our Website: http://www.httrack.com
*/
/* ------------------------------------------------------------ */
/* File: in-progress display rows, shared by httrack and htsserver .h */
/* Author: Xavier Roche */
/* ------------------------------------------------------------ */
#ifndef HTS_STATS_DEFH
#define HTS_STATS_DEFH
#include "htsglobal.h"
/* One row of the "in progress" display, shared so the CLI (httrack) and the
web GUI (htsserver) cannot drift apart: each fills its own array. */
#define NStatsBuffer 14
#ifndef HTS_DEF_FWSTRUCT_t_StatsBuffer
#define HTS_DEF_FWSTRUCT_t_StatsBuffer
typedef struct t_StatsBuffer t_StatsBuffer;
#endif
struct t_StatsBuffer {
char name[1024];
char file[1024];
char state[288]; // a short label plus back->info[256]
char BIGSTK url_sav[HTS_URLMAXSIZE * 2]; // pour cancel
char BIGSTK url_adr[HTS_URLMAXSIZE * 2];
char BIGSTK url_fil[HTS_URLMAXSIZE * 2];
LLint size;
LLint sizetot;
int offset;
//
int back;
//
int actived; // pour disabled
};
#endif

View File

@@ -276,6 +276,21 @@ int ident_url_relatif(const char *lien, const char *origin_adr,
return ok;
}
/* Bounded substring search: bodies and archive records carry NUL bytes, so
strstr() would stop at the first one. */
const char *hts_memstr(const char *hay, size_t haylen, const char *needle,
size_t nlen) {
size_t i;
if (nlen == 0 || haylen < nlen)
return NULL;
for (i = 0; i + nlen <= haylen; i++) {
if (hay[i] == *needle && memcmp(hay + i, needle, nlen) == 0)
return hay + i;
}
return NULL;
}
// créer dans s, à partir du chemin courant curr_fil, le lien vers link (absolu)
// un ident_url_relatif a déja été fait avant, pour que link ne soit pas un chemin relatif
int lienrelatif(char *s, size_t ssize, const char *link, const char *curr_fil) {
@@ -308,8 +323,10 @@ int lienrelatif(char *s, size_t ssize, const char *link, const char *curr_fil) {
// copy only the current path
curr = _curr;
strlcpybuff(curr, curr_fil, sizeof(_curr));
if ((a = strchr(curr, '?')) == NULL) // couper au ? (params)
a = curr + strlen(curr) - 1; // pas de params: aller à la fin
if ((a = strchr(curr, '?')) == NULL) { // cut at the ? (query parameters)
// an empty path has no last character: curr-1 would read before the buffer
a = curr[0] != '\0' ? curr + strlen(curr) - 1 : curr;
}
while((*a != '/') && (a > curr))
a--; // chercher dernier / du chemin courant
if (*a == '/')

View File

@@ -61,6 +61,11 @@ typedef struct lien_adrfilsave lien_adrfilsave;
int ident_url_relatif(const char *lien, const char *origin_adr,
const char *origin_fil,
lien_adrfil* const adrfil);
/* Bounded substring search over data that may hold NUL bytes; NULL if absent.
*/
const char *hts_memstr(const char *hay, size_t haylen, const char *needle,
size_t nlen);
int lienrelatif(char *s, size_t ssize, const char *link, const char *curr);
int link_has_authority(const char *lien);
int link_has_authorization(const char *lien);

View File

@@ -35,26 +35,10 @@ Please visit our Website: http://www.httrack.com
#include "htsglobal.h"
#include "htscore.h"
#include "htsstats.h"
#define NStatsBuffer 14
#define MAX_LEN_INPROGRESS 40
typedef struct t_StatsBuffer {
char name[1024];
char file[1024];
char state[288]; // a short label plus back->info[256]
char url_sav[HTS_URLMAXSIZE * 2]; // pour cancel
char url_adr[HTS_URLMAXSIZE * 2];
char url_fil[HTS_URLMAXSIZE * 2];
LLint size;
LLint sizetot;
int offset;
//
int back;
//
int actived; // pour disabled
} t_StatsBuffer;
typedef struct t_InpInfo {
int ask_refresh;
int refresh;

View File

@@ -151,6 +151,23 @@ static hts_boolean is_embed_pair(const htspair_t *table, const char *tag,
return HTS_FALSE;
}
/* The engine's robots.txt verdict for (adr,fil). Under HTS_ROBOTS_SOMETIMES an
explicit filter acceptance overrides the ban, which is why the filter outcome
is an input; the sitemap fetcher asks the same question outside the wizard.
*/
hts_boolean hts_robots_forbids(httrackp *opt, const char *adr, const char *fil,
hts_boolean filters_decided,
hts_boolean filters_refused) {
if (!opt->robots || opt->robotsptr == NULL)
return HTS_FALSE;
if (checkrobots((robots_wizard *) opt->robotsptr, adr, fil) != -1)
return HTS_FALSE;
if (filters_decided && !filters_refused &&
opt->robots == HTS_ROBOTS_SOMETIMES)
return HTS_FALSE;
return HTS_TRUE;
}
static int hts_acceptlink_(httrackp * opt, int ptr,
const char *adr, const char *fil, const char *tag,
const char *attribute, int *set_prio_to,
@@ -576,30 +593,26 @@ static int hts_acceptlink_(httrackp * opt, int ptr,
}
}
// vérifier robots.txt
if (opt->robots) {
int r = checkrobots(_ROBOTS, adr, fil);
if (r == -1) { // interdiction
if (opt->robots && checkrobots(_ROBOTS, adr, fil) == -1) {
#if DEBUG_ROBOTS
printf("robots.txt forbidden: %s%s\n", adr, fil);
printf("robots.txt forbidden: %s%s\n", adr, fil);
#endif
// question résolue, par les filtres, et mode robot non strict
if ((!question) && (filters_answer) &&
(opt->robots == HTS_ROBOTS_SOMETIMES) && (forbidden_url != 1)) {
r = 0; // annuler interdiction des robots
if (!forbidden_url) {
hts_log_print(opt, LOG_DEBUG,
"Warning link followed against robots.txt: link %s at %s%s",
l, adr, fil);
}
}
if (r == -1) { // interdire
forbidden_url = 1;
question = 0;
hts_log_print(opt, LOG_DEBUG,
"(robots.txt) forbidden link: link %s at %s%s", l, adr,
fil);
if (!hts_robots_forbids(opt, adr, fil,
(!question && filters_answer) ? HTS_TRUE
: HTS_FALSE,
(forbidden_url == 1) ? HTS_TRUE : HTS_FALSE)) {
if (!forbidden_url) {
hts_log_print(
opt, LOG_DEBUG,
"Warning link followed against robots.txt: link %s at %s%s", l,
adr, fil);
}
} else {
forbidden_url = 1;
question = 0;
hts_log_print(opt, LOG_DEBUG,
"(robots.txt) forbidden link: link %s at %s%s", l, adr,
fil);
}
}

View File

@@ -49,6 +49,13 @@ typedef struct httrackp httrackp;
typedef struct lien_url lien_url;
#endif
/* The engine's robots.txt verdict for (adr,fil): HTS_TRUE when the fetch is
forbidden. `filters_decided`/`filters_refused` carry the filter outcome,
which overrides a ban under -s1 (HTS_ROBOTS_SOMETIMES). */
hts_boolean hts_robots_forbids(httrackp *opt, const char *adr, const char *fil,
hts_boolean filters_decided,
hts_boolean filters_refused);
int hts_acceptlink(httrackp * opt, int ptr,
const char *adr, const char *fil,
const char *tag, const char *attribute,

View File

@@ -230,8 +230,7 @@ static hts_boolean vt_size_refresh(void) {
*/
#define STYLE_STATVALUES VT_BOLD
#define STYLE_STATTEXT VT_UNBOLD
#define STYLE_STATRESET VT_UNBOLD
#define NStatsBuffer 14
#define STYLE_STATRESET VT_UNBOLD
/* Rows the stats block and "Current job" take above the in-progress list. */
#define NStatsHeaderRows 7

View File

@@ -36,26 +36,7 @@ Please visit our Website: http://www.httrack.com
#include "htsglobal.h"
#include "htscore.h"
#include "htssafe.h"
#ifndef HTS_DEF_FWSTRUCT_t_StatsBuffer
#define HTS_DEF_FWSTRUCT_t_StatsBuffer
typedef struct t_StatsBuffer t_StatsBuffer;
#endif
struct t_StatsBuffer {
char name[1024];
char file[1024];
char state[288]; // a short label plus back->info[256]
char BIGSTK url_sav[HTS_URLMAXSIZE * 2]; // pour cancel
char BIGSTK url_adr[HTS_URLMAXSIZE * 2];
char BIGSTK url_fil[HTS_URLMAXSIZE * 2];
LLint size;
LLint sizetot;
int offset;
//
int back;
//
int actived; // pour disabled
};
#include "htsstats.h"
#ifndef HTS_DEF_FWSTRUCT_t_InpInfo
#define HTS_DEF_FWSTRUCT_t_InpInfo

View File

@@ -138,6 +138,7 @@
<ClCompile Include="htswrap.c" />
<ClCompile Include="htszlib.c" />
<ClCompile Include="htswarc.c" />
<ClCompile Include="htssitemap.c" />
<ClCompile Include="md5.c" />
<ClCompile Include="minizip\ioapi.c" />
<ClCompile Include="minizip\iowin32.c" />

View File

@@ -2380,8 +2380,18 @@ static int PT_SaveCache__Arc_Fun(void *arg, const char *url, PT_Element element)
PT_SaveCache__Arc_t *st = (PT_SaveCache__Arc_t *) arg;
FILE *const fp = st->fp;
struct tm *tm = convert_time_rfc822(&st->buff, element->lastmodified);
struct tm unknown_date;
int size_headers;
/* a cached entry with no parseable Last-Modified must not take the writer
down; the epoch is the conventional "date unknown" */
if (tm == NULL) {
memset(&unknown_date, 0, sizeof(unknown_date));
unknown_date.tm_year = 70;
unknown_date.tm_mday = 1;
tm = &unknown_date;
}
sprintf(st->headers,
"HTTP/1.0 %d %s"
"\r\n"

View File

@@ -0,0 +1,10 @@
#!/bin/bash
#
# Sitemap parser self-test: <loc> extraction, entity decoding, URL and length
# rejections, the URL cap, gzip framing and robots.txt Sitemap: records.
set -euo pipefail
out=$(httrack -O /dev/null '-#test=sitemap')
echo "$out"
test "$out" = "sitemap self-test OK"

View File

@@ -146,11 +146,13 @@ sid=$(scrape_sid "${port}")
resp=$(post_redirect '/foo
X-Injected: pwned' "${port}" "${sid}") || fail "no response to a CRLF redirect value"
stop
printf '%s' "${resp}" | grep -qi 'X-Injected' &&
# Here-string, not a pipe: grep -q SIGPIPEs the producer on the first match, and
# pipefail then turns the hit into a silent pass.
grep -qi 'X-Injected' <<<"${resp}" &&
fail "CRLF in the redirect value reached the response"
test "$(printf '%s' "${resp}" | grep -ci '^Location:')" -eq 0 ||
test "$(grep -ci '^Location:' <<<"${resp}")" -eq 0 ||
fail "CRLF redirect still emitted a Location header"
test "$(printf '%s' "${resp}" | grep -c '^HTTP/1\.')" -eq 1 ||
test "$(grep -c '^HTTP/1\.' <<<"${resp}")" -eq 1 ||
fail "CRLF redirect did not yield exactly one status line"
# Loopback unless asked otherwise. Assert the socket, not the announcement.

View File

@@ -103,9 +103,10 @@ sid=$(scrape_sid "${port}")
test "${#sid}" -eq 32 || fail "did not scrape a 32-hex sid from the page (got '${sid}')"
# Accept: the legitimate flow still works. "redirect" is the cheapest field
# with a reply that is visible in the headers.
# with a reply that is visible in the headers. Matched from a here-string, not a
# pipe: grep -q SIGPIPEs the producer, and pipefail then buries the verdict.
resp=$(request "${port}" "sid=${sid}&redirect=/accepted")
printf '%s' "${resp}" | grep -q '^Location: /accepted' ||
grep -q '^Location: /accepted' <<<"${resp}" ||
fail "a body carrying the correct sid was refused"
# Refuse: missing and wrong. A missing one is the regression under test — the
@@ -114,12 +115,12 @@ printf '%s' "${resp}" | grep -q '^Location: /accepted' ||
for bad in "redirect=/nosid" "sid=&redirect=/empty" \
"sid=00000000000000000000000000000000&redirect=/wrong"; do
resp=$(request "${port}" "${bad}")
printf '%s' "${resp}" | grep -q '^Location:' &&
grep -q '^Location:' <<<"${resp}" &&
fail "body accepted without a valid sid: ${bad}"
# A refusal has to be a well-formed reply, not a headerless fragment: the
# release build used to emit only Content-length, which reads as a protocol
# error to any client and hides the reason.
printf '%s' "${resp}" | grep -q '^HTTP/1\.0 403 ' ||
grep -q '^HTTP/1\.0 403 ' <<<"${resp}" ||
fail "refusal was not a 403: ${bad}"
done
@@ -127,16 +128,23 @@ done
# applied to one global key store that later requests render from, and the
# command dispatcher reads it from there. Probe the store itself — step3
# interpolates ${projname} into its title — rather than the refused reply.
title() { request "$1" "" step3; }
store_page() {
# a dead probe or a non-page reply must not read as "the store was not written"
page=$(request "$1" "" step3) || fail "the key store probe failed"
grep -q '^HTTP/1\.0 200 ' <<<"${page}" ||
fail "the key store probe did not answer: '$(head -1 <<<"${page}")'"
}
request "${port}" "projname=UNAUTHWRITE" >/dev/null 2>&1 || true
title "${port}" | grep -q 'UNAUTHWRITE' &&
store_page "${port}"
grep -q 'UNAUTHWRITE' <<<"${page}" &&
fail "a body without a valid sid was written to the key store"
# Paired accept case: without it the assertion above passes even if the server
# simply ignores every body, which would prove nothing.
request "${port}" "sid=${sid}&projname=AUTHWRITE" >/dev/null 2>&1 || true
title "${port}" | grep -q 'AUTHWRITE' ||
store_page "${port}"
grep -q 'AUTHWRITE' <<<"${page}" ||
fail "a body carrying the correct sid was not written to the key store"
echo "PASS"

View File

@@ -90,6 +90,15 @@ sys.stdout.write(out.decode("latin-1"))' "$1" "$2" "$3"
get() { request "$1" "$2" ""; }
post() { request "$1" "" "$2"; }
# GET $2 into ${reply}, requiring status $3. Captured, not piped: a dead request
# reads as marker-absent, and so do a truncated body and a 302 to the file.
reply=
fetch() {
reply=$(get "$1" "$2") || fail "the request for $2 failed"
grep -q "^HTTP/1\.0 $3 " <<<"${reply}" ||
fail "$2: wanted a $3 reply, got '$(head -1 <<<"${reply}")'"
}
url=$(start)
port=$(portof "${url}")
srv=$(srvpid)
@@ -109,13 +118,14 @@ echo "LOGMARKER" >"${base}/proj/hts-log.txt"
# with a component over NAME_MAX (mkdir refuses it whatever the uid). Must precede
# any successful save: commandEnd then swaps the error page for the finished one.
get "${port}" /server/style.css >/dev/null # leaves a path behind in fsfile
post "${port}" "sid=${sid}&command=httrack&command_do=save&winprofile=x&path=${base}/$(printf '%0300d' 0)&projname=p" |
grep -q '^Location: /server/error.html' ||
saved=$(post "${port}" "sid=${sid}&command=httrack&command_do=save&winprofile=x&path=${base}/$(printf '%0300d' 0)&projname=p")
grep -q '^Location: /server/error.html' <<<"${saved}" ||
fail "the refused save did not redirect to the error page"
# No project yet, so no root to serve from.
post "${port}" "sid=${sid}&projpath=${base}/" >/dev/null
get "${port}" /website/secret.txt | grep -q SECRETMARKER &&
fetch "${port}" /website/secret.txt 404
grep -q SECRETMARKER <<<"${reply}" &&
fail "a posted projpath served a file outside any project"
# Positive control: step4's "save settings" flow registers the project without
@@ -126,24 +136,29 @@ body="${body}&path=${base}&projname=proj&projpath=${base}/proj/"
post "${port}" "${body}" >/dev/null
test -f "${base}/proj/hts-cache/winprofile.ini" ||
fail "the project was not registered: $(cat "${log}")"
get "${port}" /website/hts-log.txt | grep -q LOGMARKER ||
fetch "${port}" /website/hts-log.txt 200
grep -q LOGMARKER <<<"${reply}" ||
fail "the registered project's mirror is not served"
# Same request with the root repointed: the project is legitimate, projpath is
# not what decides where the bytes come from.
post "${port}" "sid=${sid}&projpath=${base}/" >/dev/null
get "${port}" /website/secret.txt | grep -q SECRETMARKER &&
fetch "${port}" /website/secret.txt 404
grep -q SECRETMARKER <<<"${reply}" &&
fail "a posted projpath repointed the served root"
post "${port}" "sid=${sid}&projpath=/etc/" >/dev/null
get "${port}" /website/passwd | grep -q '^root:' &&
fetch "${port}" /website/passwd 404
grep -q '^root:' <<<"${reply}" &&
fail "a posted projpath read an arbitrary system file"
# A ".." anywhere in the recorded root would escape the mirror on every later
# request, so the save must be refused and the previous root kept.
post "${port}" "sid=${sid}&command=httrack&command_do=save&winprofile=x&path=${base}/proj&projname=.." >/dev/null
get "${port}" /website/secret.txt | grep -q SECRETMARKER &&
fetch "${port}" /website/secret.txt 404
grep -q SECRETMARKER <<<"${reply}" &&
fail "a '..' in the saved project path escaped the mirror root"
get "${port}" /website/hts-log.txt | grep -q LOGMARKER ||
fetch "${port}" /website/hts-log.txt 200
grep -q LOGMARKER <<<"${reply}" ||
fail "rejecting the '..' root also lost the previous one"
# The root is now the server's own, but it is still built from two posted
@@ -162,7 +177,7 @@ test -f "${fspath}/hts-cache/winprofile.ini" ||
fail "the long-path project was not registered: $(cat "${log}")"
get "${port}" "/website/${longurl}" >/dev/null 2>&1 || true
alive "${srv}" || fail "an over-long project path crashed the server: $(cat "${log}")"
get "${port}" /server/index.html | grep -q '200 OK' ||
fail "the server stopped answering after the over-long project path"
# Not just alive: still answering.
fetch "${port}" /server/index.html 200
echo "PASS"

View File

@@ -0,0 +1,51 @@
#!/bin/bash
# A cached entry whose Last-Modified is missing or unparseable must still
# convert: the ARC writer used to dereference the parsed date unchecked.
set -euo pipefail
dir=$(mktemp -d)
trap 'rm -rf "$dir"' EXIT
# $1 Last-Modified value (empty = omit the header), $2 expected archive date,
# $3 label. The date is asserted exactly: a guard that fires unconditionally
# would clobber the valid case, and one that skips the year/day would emit a
# month and day of 00.
run_case() {
lastmod=$1
wantdate=$2
what=$3
{
printf 'HTTP/1.1 200 OK\r\n'
printf 'Content-Type: text/html\r\n'
test -z "$lastmod" || printf 'Last-Modified: %s\r\n' "$lastmod"
printf 'Content-Length: 5\r\n\r\n'
} >"$dir/hdr"
printf 'hello' >"$dir/body"
alen=$(($(wc -c <"$dir/hdr") + $(wc -c <"$dir/body")))
{
printf 'filedesc://t.arc 0.0.0.0 20250101000000 text/plain 200 - - 0 t.arc 9\n'
printf '2 0 test\n'
printf '\n\n'
printf 'http://example.com/p.html 0.0.0.0 20250101000000 text/html 200 - - 0 t.arc %d\n' "$alen"
cat "$dir/hdr" "$dir/body"
} >"$dir/in.arc"
proxytrack --convert "$dir/out.arc" "$dir/in.arc" >/dev/null 2>&1 || {
echo "FAIL: proxytrack crashed on $what" >&2
exit 1
}
grep -aq "^http://example.com/p.html 0.0.0.0 ${wantdate} " "$dir/out.arc" || {
echo "FAIL: $what: expected archive date ${wantdate}, got:" >&2
grep -a '^http://example.com/' "$dir/out.arc" >&2 || echo "(entry dropped)" >&2
exit 1
}
}
# the epoch stands in for "date unknown"
run_case '' 19700101000000 'a missing Last-Modified'
run_case 'not a date at all' 19700101000000 'an unparseable Last-Modified'
run_case '0' 19700101000000 'a bare 0 Last-Modified'
# a parseable date must still come through untouched
run_case 'Sun, 06 Nov 1994 08:49:37 GMT' 19941106084937 'a valid Last-Modified'

134
tests/89_local-sitemap.test Normal file
View File

@@ -0,0 +1,134 @@
#!/bin/bash
#
# --sitemap seeds the crawl from robots.txt -> sitemapindex -> gzipped urlset.
# start.html links to nothing, so orphan*.html can only arrive through the
# sitemap; deep1.html proves the seeds keep a full depth budget under -r2, and
# the off-host page and child sitemap must both be refused.
set -eu
: "${top_srcdir:=..}"
crawl() { bash "$top_srcdir/tests/local-crawl.sh" "$@"; }
# robots.txt Sitemap: -> index -> .xml.gz, seeds behaving like -r2 seeds. The
# log assertions pin which route was taken: the two documents are served from
# both /sitemapdir/index.xml and the well-known /sitemap.xml, so "some sitemap
# was read" would pass either way. --not-found pins ingestion-only: the sitemap
# documents feed the crawl but never land in the mirror.
# --rerun also walks the update path over a mirror holding sitemap seeds.
crawl --errors 0 --rerun \
--found 'sitemapdir/orphan1.html' \
--found 'sitemapdir/orphan2.html' \
--found 'sitemapdir/deep1.html' \
--not-found 'sitemapdir/index.xml' \
--not-found 'sitemapdir/pages.xml.gz' \
--log-found '2 of 3 URL\(s\) added from 127\.0\.0\.1:[0-9]+/sitemapdir/pages\.xml\.gz' \
--log-found '1 of 2 child sitemap\(s\) listed by 127\.0\.0\.1:[0-9]+/sitemapdir/index\.xml' \
--log-found 'ignoring off-host child sitemap' \
--log-not-found '/sitemap\.xml' \
httrack 'BASEURL/sitemapdir/start.html' --sitemap -r2
# Negative control: without the option nothing but the start page is reached.
crawl --errors 0 \
--found 'sitemapdir/start.html' \
--not-found 'sitemapdir/orphan1.html' \
--not-found 'sitemapdir/orphan2.html' \
httrack 'BASEURL/sitemapdir/start.html' -r2
# Sitemap URLs are not a filter bypass: a -*orphan2* rule still rejects one.
crawl --errors 0 \
--found 'sitemapdir/orphan1.html' \
--not-found 'sitemapdir/orphan2.html' \
httrack 'BASEURL/sitemapdir/start.html' --sitemap -r2 '-*orphan2*'
# A robots.txt naming no sitemap falls back to the well-known /sitemap.xml.
# The test server drops its Sitemap: record for this User-Agent.
crawl --errors 0 \
--found 'sitemapdir/orphan1.html' \
--log-found 'listed by 127\.0\.0\.1:[0-9]+/sitemap\.xml' \
--log-not-found 'sitemapdir/index\.xml' \
httrack 'BASEURL/sitemapdir/start.html' --sitemap -r2 -F 'nositemap-agent'
# --sitemap-url alone reads the named document and probes nothing.
crawl --errors 0 \
--found 'sitemapdir/orphan1.html' \
--found 'sitemapdir/orphan2.html' \
--log-not-found 'sitemapdir/index\.xml' \
--log-not-found '/sitemap\.xml' \
httrack 'BASEURL/sitemapdir/start.html' -r2 \
--sitemap-url 'BASEURL/sitemapdir/pages.xml.gz'
# The two options are additive: an unfetchable --sitemap-url is parsed as an
# empty document and does not stop --sitemap from seeding the crawl.
crawl --errors-content 1 \
--found 'sitemapdir/orphan1.html' \
--log-found 'Sitemap: 0 of 0 URL\(s\) added' \
httrack 'BASEURL/sitemapdir/start.html' --sitemap -r2 \
--sitemap-url 'BASEURL/sitemapdir/missing.xml'
# The sitemapindex nesting cap stops the chain before its deepest urlset. The
# positive control is the page one level inside the cap: without it, a cap
# mutated to 0 (nothing followed) would pass this just as well.
crawl --errors 0 \
--found 'sitemapdir/cap3.html' \
--not-found 'sitemapdir/cap4.html' \
--log-found 'Sitemap: cap reached' \
httrack 'BASEURL/sitemapdir/start.html' -r2 \
--sitemap-url 'BASEURL/sitemapdir/chain0.xml'
# A site cannot widen a subtree crawl by putting its sitemap at the root: the
# page below the start directory is seeded, the one above it is refused.
crawl --errors 0 \
--found 'deep/dir/below.html' \
--not-found 'elsewhere/updir.html' \
httrack 'BASEURL/deep/dir/start.html' --sitemap -r3 -F 'scopesitemap-agent'
# How far a sitemap fetch is gated depends on who asked for it, so the three
# cases below have to differ: one "robots is honoured" assertion would hide a
# regression in either direction. The harness disables robots, hence -s2 here.
# We guessed /sitemap.xml, so a Disallow on it wins.
crawl --errors 0 \
--not-found 'sitemapdir/orphan1.html' \
--log-found 'Sitemap: robots.txt forbids' \
httrack 'BASEURL/sitemapdir/start.html' --sitemap -r2 -s2 \
-F 'denysitemap-agent'
# The site declared it through a Sitemap: line, which the same Disallow does
# not retract.
crawl --errors 0 \
--found 'sitemapdir/orphan1.html' \
--log-not-found 'Sitemap: robots.txt forbids' \
httrack 'BASEURL/sitemapdir/start.html' --sitemap -r2 -s2 \
-F 'denydeclared-agent'
# The user named it, which is the same intent as a start URL: not refused
# either, even though robots.txt forbids that exact path.
crawl --errors 0 \
--found 'sitemapdir/orphan1.html' \
--log-not-found 'Sitemap: robots.txt forbids' \
httrack 'BASEURL/sitemapdir/start.html' -r2 -s2 -F 'denysitemap-agent' \
--sitemap-url 'BASEURL/sitemap.xml'
# A child sitemap is a fetch like any other: a filter rejecting it stops the
# whole subtree, so the page only it lists is never reached.
crawl --errors 0 \
--not-found 'sitemapdir/gated.html' \
--log-found 'Sitemap: filter rule #[0-9]+ refuses' \
httrack 'BASEURL/sitemapdir/start.html' -r2 '-*filtered.xml*' \
--sitemap-url 'BASEURL/sitemapdir/gatedindex.xml'
# Control: without that filter the same chain does reach the page.
crawl --errors 0 \
--found 'sitemapdir/gated.html' \
httrack 'BASEURL/sitemapdir/start.html' -r2 \
--sitemap-url 'BASEURL/sitemapdir/gatedindex.xml'
# A redirected sitemap is still ingested: the marking follows the 301.
crawl --errors 0 \
--found 'sitemapdir/orphan1.html' \
--log-found 'moved\.xml redirects to .*pages\.xml\.gz' \
--log-found 'URL\(s\) added from 127\.0\.0\.1:[0-9]+/sitemapdir/pages\.xml\.gz' \
httrack 'BASEURL/sitemapdir/start.html' -r2 \
--sitemap-url 'BASEURL/sitemapdir/moved.xml'

View File

@@ -0,0 +1,196 @@
#!/bin/bash
#
# An unchecked box posts nothing and the stored "1" survives: hence a hidden
# companion per box, plus a disabling flag for the default-on ones.
set -euo pipefail
testdir=$(cd "$(dirname "$0")" && pwd)
distdir=${top_srcdir:-$(cd "${testdir}/.." && pwd)}
distdir=$(cd "${distdir}" && pwd)
# shellcheck source=tests/testlib.sh
. "${testdir}/testlib.sh"
fail() {
echo "FAIL: $*" >&2
exit 1
}
command -v htsserver >/dev/null || fail "no htsserver in PATH"
python=$(find_python) || {
echo "python3 not found; skipping" >&2
exit 77
}
work=$(mktemp -d "${TMPDIR:-/tmp}/webhttrack_checkbox.XXXXXX") || fail "no tmpdir"
srvlog=$(mktemp)
srv=
srvpid=
cleanup() {
# htsserver keeps SIGTERM ignored across its exec, so only -9 reaps it.
test -z "${srvpid}" || kill -9 "${srvpid}" 2>/dev/null || true
test -z "${srv}" || kill -9 "${srv}" 2>/dev/null || true
wait "${srv}" 2>/dev/null || true # absorb bash's async "Killed" notice
rm -rf "${work}" "${srvlog}"
}
trap cleanup EXIT HUP INT QUIT PIPE TERM
# An isolated HOME keeps a stray ~/.httrack.ini out of the served settings.
sport=$("${python}" -c 'import socket
s = socket.socket()
s.bind(("127.0.0.1", 0))
print(s.getsockname()[1])
s.close()')
(
trap '' TERM TTOU
export HOME="${work}"
exec htsserver "${distdir}/" --port "${sport}" >"${srvlog}" 2>&1
) &
srv=$!
for _ in $(seq 1 40); do
url=$(sed -n 's/^URL=//p' "${srvlog}") && test -n "${url}" && break
kill -0 "${srv}" 2>/dev/null || break
sleep 0.25
done
test -n "${url:-}" || fail "htsserver did not start: $(cat "${srvlog}")"
srvpid=$(sed -n 's/^PID=//p' "${srvlog}") # absent on Windows
"${python}" - "${url}" "${distdir}" <<'PY' || fail "checkbox clearing is broken (see above)"
import glob, os, re, sys, urllib.parse, urllib.request
url, srcdir = sys.argv[1].rstrip("/"), sys.argv[2]
rc = 0
def check(ok, what):
global rc
print(("ok: " if ok else "FAIL: ") + what)
if not ok:
rc = 1
# cache/cache2 are left to the separate rework of the dead "Cache" seed.
skip = {"cache", "cache2"}
page_of = {}
boxes = 0
for path in sorted(glob.glob(os.path.join(srcdir, "html", "server", "option*.html"))):
page = open(path, "rb").read().decode("latin-1")
cleared = dict((m.group(1), m.start()) for m in
re.finditer(r'<input type="hidden" name="([^"]+)" value="">', page))
for m in re.finditer(r'<input type="checkbox" name="([^"]+)"', page):
boxes += 1
if m.group(1) in skip:
continue
page_of[m.group(1)] = os.path.basename(path)
at = cleared.get(m.group(1))
check(at is not None and at < m.start(), "%s: %s is cleared before it is drawn"
% (os.path.basename(path), m.group(1)))
# Control: a regex that stopped matching would leave nothing to assert on.
check(boxes >= 25, "the option pages were scanned (%d checkboxes)" % boxes)
sid = None
def get(path):
# The UI is served ISO-8859-1, so decode, do not assume UTF-8.
return urllib.request.urlopen(url + path, timeout=20).read().decode("latin-1")
def post(fields):
global sid
if sid is None:
m = re.search(r'name="sid" value="([0-9a-f]+)"', get("/server/index.html"))
if m is None:
sys.exit("no session id in server/index.html")
sid = m.group(1)
body = "&".join("%s=%s" % (k, urllib.parse.quote(v)) for k, v in
[("sid", sid)] + fields)
req = urllib.request.Request(url + "/server/step4.html",
data=body.encode("latin-1"), method="POST")
return urllib.request.urlopen(req, timeout=20).read().decode("latin-1")
def textarea(page, name):
m = re.search(r'<textarea name="%s".*?>(.*?)</textarea>' % name, page, re.S)
if m is None:
sys.exit("no %s textarea in the rendered step4.html" % name)
return m.group(1)
def checked(page, name):
return re.search(r'name="%s"[ \t]*checked' % name, get("/server/" + page)) is not None
# What each state puts on the command line (None: nothing) and its profile key.
# Only a value-taking alias can carry a disabling flag: -I and -%I are "single"
# (htsalias.c), so a --index=0 would read back as the bare enabling flag.
BOXES = [
# field profile key set cleared
("parseall", "ParseAll", "--near", None),
("link", "Near", "--test", None),
("testall", "Test", "--extended-parsing", None),
# htmlfirst emits --priority=7, which the scan-priority list also emits.
("htmlfirst", "HTMLFirst", None, None),
("errpage", "NoErrorPages", "--generate-errors=0", None),
("external", "NoExternalPages", "--replace-external", None),
("hidepwd", "NoPwdInPages", "--disable-passwords", None),
("hidequery", "NoQueryStrings", "--include-query-string=0", None),
("nopurge", "NoPurgeOldFiles", "--purge-old=0", None),
("windebug", None, "--debug-headers", None),
("ka", "KeepAlive", "--keep-alive", None),
("remt", "RemoveTimeout", "--host-control=1", None),
("rems", "RemoveRateout", "--host-control=2", None),
("cookies", "Cookies", None, "--cookies=0"),
("parsejava", "ParseJava", None, "--parse-java=0"),
("updhack", "UpdateHack", "--updatehack", None),
("urlhack", "URLHack", "--urlhack", None),
("keepwww", "KeepWww", "--keep-www-prefix", None),
("keepslashes", "KeepSlashes", "--keep-double-slashes", None),
("keepqueryorder", "KeepQueryOrder", "--keep-query-order", None),
("toler", "TolerantRequests", "--tolerant", None),
("http10", "HTTP10", "--http-10", None),
("sitemap", "Sitemap", "--sitemap", None),
("warc", "Warc", "--warc", None),
("norecatch", "NoRecatch", "--do-not-recatch", None),
("logf", "Log", "--single-log", None),
("index", "Index", None, None),
("index2", "WordIndex", "--search-index", None),
("ftpprox", "UseHTTPProxyForFTP", "--httpproxy-ftp", None),
]
for field, key, when_set, when_clear in BOXES:
for value, want, unwanted, stored in (("on", when_set, when_clear, "1"),
("", when_clear, when_set, "0")):
rendered = post([(field, value)])
cmd = textarea(rendered, "command").split()
label = "%s %s" % (field, "set" if value else "cleared")
if want:
check(want in cmd, "%s emits %s" % (label, want))
if unwanted:
check(unwanted not in cmd, "%s drops %s" % (label, unwanted))
if key:
# The Windows GUI reads the same option out of the profile.
check("%s=%s" % (key, stored) in textarea(rendered, "winprofile").splitlines(),
"%s writes %s=%s" % (label, key, stored))
check(checked(page_of[field], field) == bool(value),
"%s draws %s box" % (label, "a ticked" if value else "an empty"))
missing = sorted(set(page_of) - set(b[0] for b in BOXES))
check(not missing, "every checkbox is exercised (missing %s)" % missing)
# A browser posts the companion and the ticked box under one name, so the fix
# rests on the body loop keeping the last value; both orders pin that down.
cmd = textarea(post([("cookies", ""), ("cookies", "on")]), "command").split()
check("--cookies=0" not in cmd, "cookies=&cookies=on keeps cookies set")
check(checked("option8.html", "cookies"), "cookies=&cookies=on draws a ticked box")
cmd = textarea(post([("cookies", "on"), ("cookies", "")]), "command").split()
check("--cookies=0" in cmd, "cookies=on&cookies= clears cookies")
check(not checked("option8.html", "cookies"), "cookies=on&cookies= draws an empty box")
sys.exit(rc)
PY
# A leaked htsserver wedges the parallel harness behind a green log.
cleanup
! kill -0 "${srv}" 2>/dev/null || fail "htsserver ${srv} survived"
echo "PASS"

View File

@@ -90,6 +90,7 @@ TESTS = \
01_engine-xfread.test \
01_zlib-acceptencoding.test \
01_zlib-warc.test \
01_zlib-sitemap.test \
01_zlib-warc-cdx.test \
01_zlib-warc-wacz.test \
01_zlib-contentcodings.test \
@@ -183,6 +184,9 @@ TESTS = \
83_webhttrack-argescape.test \
84_webhttrack-mirror-verbatim.test \
85_webhttrack-projpath.test \
86_local-proxytrack-cache-longfields.test
86_local-proxytrack-cache-longfields.test \
87_local-proxytrack-nodate.test \
89_local-sitemap.test \
90_webhttrack-checkbox-clear.test
CLEANFILES = check-network_sh.cache

View File

@@ -18,6 +18,7 @@ import base64
import gzip
import hashlib
import os
import re
import sys
import time
from http.server import SimpleHTTPRequestHandler, ThreadingHTTPServer
@@ -542,8 +543,36 @@ class Handler(SimpleHTTPRequestHandler):
return self.fail_cookie(name)
self.send_html("\tThis is the secret.")
# A User-Agent carrying NO_SITEMAP_UA gets a robots.txt with no Sitemap:
# record, so a test can drive the /sitemap.xml fallback instead.
NO_SITEMAP_UA = "nositemap"
# ... and one that additionally Disallows the well-known location, so the
# fallback has to be refused by the rules this very body carries.
DENY_SITEMAP_UA = "denysitemap"
# ... the same Disallow, but with the sitemap also declared: the
# declaration is the site inviting the fetch and must win.
DENY_DECLARED_UA = "denydeclared"
# ... and one that points the sitemap at the site root, to check a subtree
# crawl is not widened by where the site chooses to put its sitemap.
SCOPE_SITEMAP_UA = "scopesitemap"
def route_robots(self):
body = b"User-agent: *\nDisallow:\n"
# The Sitemap: record is group-independent; only --sitemap acts on it.
ua = self.headers.get("User-Agent") or ""
host = self.headers.get("Host")
body = "User-agent: *\nDisallow:\n"
if self.DENY_DECLARED_UA in ua:
body = (
"User-agent: *\nDisallow: /sitemap.xml\n"
f"Sitemap: http://{host}/sitemap.xml\n"
)
elif self.DENY_SITEMAP_UA in ua:
body = "User-agent: *\nDisallow: /sitemap.xml\n"
elif self.SCOPE_SITEMAP_UA in ua:
body += f"Sitemap: http://{host}/scopesitemap.xml\n"
elif self.NO_SITEMAP_UA not in ua:
body += f"Sitemap: http://{host}/sitemapdir/index.xml\n"
body = body.encode()
self.send_response(200)
self.send_header("Content-Type", "text/plain")
self.send_header("Content-Length", str(len(body)))
@@ -551,6 +580,132 @@ class Handler(SimpleHTTPRequestHandler):
if self.command != "HEAD":
self.wfile.write(body)
# --- sitemap ingestion (issue #712) ------------------------------------
# start.html links to nothing, so orphan*.html are reachable only through
# the sitemap. deep1.html proves the seeds keep a full depth budget; the
# off-host page <loc> must be dropped by the travel scope, and the off-host
# child sitemap by the ingester's same-host rule. The index is served both
# from /sitemapdir/ (named by robots.txt) and from the well-known
# /sitemap.xml (the fallback).
def route_sitemap_index(self):
host = self.headers.get("Host")
self.send_raw(
'<?xml version="1.0" encoding="UTF-8"?>\n'
'<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">'
f"<sitemap><loc>http://{host}/sitemapdir/pages.xml.gz</loc></sitemap>"
"<sitemap><loc>http://sitemap-offhost.invalid/s.xml</loc></sitemap>"
"</sitemapindex>\n".encode(),
"application/xml",
)
def route_sitemap_pages(self):
host = self.headers.get("Host")
xml = (
'<?xml version="1.0" encoding="UTF-8"?>\n'
'<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">'
f"<url><loc>http://{host}/sitemapdir/orphan1.html</loc></url>"
"<url><loc>http://sitemap-offhost.invalid/x.html</loc></url>"
f"<url><loc>http://{host}/sitemapdir/orphan2.html</loc></url>"
"</urlset>\n"
).encode()
self.send_raw(gzip.compress(xml), "application/x-gzip")
def route_sitemap_start(self):
self.send_html("\tNothing links to the sitemap pages.")
def route_sitemap_orphan1(self):
self.send_html('\t<a href="deep1.html">deeper</a>')
def route_sitemap_orphan2(self):
self.send_html("\tSecond orphan.")
def route_sitemap_deep1(self):
self.send_html("\tOne level below an orphan.")
# chainN is a sitemapindex at nesting level N, listing chain(N+1) and a
# urlset capN.xml whose single page is capN.html. Levels up to
# HTS_SITEMAP_MAX_LEVEL are followed, so capN.html appears for N below the
# cap and stops appearing at it: the pair pins the boundary, which a cap
# mutated either way would break.
def route_sitemap_chain(self):
host = self.headers.get("Host")
level = int(self.path.rsplit("/", 1)[-1][len("chain") : -len(".xml")])
self.send_raw(
(
'<?xml version="1.0" encoding="UTF-8"?>\n<sitemapindex>'
f"<sitemap><loc>http://{host}/sitemapdir/chain{level + 1}.xml"
"</loc></sitemap>"
f"<sitemap><loc>http://{host}/sitemapdir/cap{level}.xml"
"</loc></sitemap></sitemapindex>\n"
).encode(),
"application/xml",
)
def route_sitemap_capset(self):
host = self.headers.get("Host")
level = self.path.rsplit("/", 1)[-1][len("cap") : -len(".xml")]
self.send_raw(
(
'<?xml version="1.0" encoding="UTF-8"?>\n<urlset>'
f"<url><loc>http://{host}/sitemapdir/cap{level}.html</loc></url>"
"</urlset>\n"
).encode(),
"application/xml",
)
def route_sitemap_cappage(self):
self.send_html("\tReached through a nested sitemapindex.")
def route_sitemap_gatedindex(self):
host = self.headers.get("Host")
self.send_raw(
'<?xml version="1.0" encoding="UTF-8"?><sitemapindex>'
f"<sitemap><loc>http://{host}/sitemapdir/filtered.xml</loc></sitemap>"
"</sitemapindex>\n".encode(),
"application/xml",
)
def route_sitemap_filtered(self):
host = self.headers.get("Host")
self.send_raw(
'<?xml version="1.0" encoding="UTF-8"?><urlset>'
f"<url><loc>http://{host}/sitemapdir/gated.html</loc></url>"
"</urlset>\n".encode(),
"application/xml",
)
def route_sitemap_gated(self):
self.send_html("\tListed only by the filtered child sitemap.")
# A moved sitemap: the marking has to follow the redirect.
# A root sitemap naming a page below the crawl's start directory and one
# above it. Only the first may be seeded when the crawl started at /deep/dir/.
def route_sitemap_scope(self):
host = self.headers.get("Host")
self.send_raw(
'<?xml version="1.0" encoding="UTF-8"?><urlset>'
f"<url><loc>http://{host}/deep/dir/below.html</loc></url>"
f"<url><loc>http://{host}/elsewhere/updir.html</loc></url>"
"</urlset>\n".encode(),
"application/xml",
)
def route_sitemap_deepstart(self):
self.send_html("\tA start page in a subdirectory, linking nothing.")
def route_sitemap_below(self):
self.send_html("\tBelow the start directory.")
def route_sitemap_updir(self):
self.send_html("\tAbove the start directory.")
def route_sitemap_moved(self):
self.send_response(301)
self.send_header("Location", "/sitemapdir/pages.xml.gz")
self.send_header("Content-Length", "0")
self.end_headers()
# --- type/extension matrix (issue #267 family) -------------------------
def send_raw(self, body, content_type, extra_headers=()):
@@ -1578,6 +1733,21 @@ class Handler(SimpleHTTPRequestHandler):
"/gated/index.php": route_gated_index,
"/gated/secret.php": route_gated_secret,
"/robots.txt": route_robots,
"/sitemapdir/index.xml": route_sitemap_index,
"/sitemap.xml": route_sitemap_index,
"/sitemapdir/pages.xml.gz": route_sitemap_pages,
"/sitemapdir/start.html": route_sitemap_start,
"/sitemapdir/orphan1.html": route_sitemap_orphan1,
"/sitemapdir/orphan2.html": route_sitemap_orphan2,
"/sitemapdir/deep1.html": route_sitemap_deep1,
"/sitemapdir/gatedindex.xml": route_sitemap_gatedindex,
"/sitemapdir/filtered.xml": route_sitemap_filtered,
"/sitemapdir/gated.html": route_sitemap_gated,
"/sitemapdir/moved.xml": route_sitemap_moved,
"/scopesitemap.xml": route_sitemap_scope,
"/deep/dir/start.html": route_sitemap_deepstart,
"/deep/dir/below.html": route_sitemap_below,
"/elsewhere/updir.html": route_sitemap_updir,
"/warcgz/index.html": route_warcgz_index,
"/warcgz/page.html": route_warcgz_page,
"/warcgz/data.bin": route_warcgz_data,
@@ -1902,6 +2072,13 @@ class Handler(SimpleHTTPRequestHandler):
return True
# Match percent-encoded paths (accented #157 route) by their decoded form.
handler = self.ROUTES.get(path) or self.ROUTES.get(unquote(path))
if handler is None:
if re.fullmatch(r"/sitemapdir/chain\d+\.xml", path):
handler = type(self).route_sitemap_chain
elif re.fullmatch(r"/sitemapdir/cap\d+\.xml", path):
handler = type(self).route_sitemap_capset
elif re.fullmatch(r"/sitemapdir/cap\d+\.html", path):
handler = type(self).route_sitemap_cappage
if handler is not None:
handler(self)
return True

View File

@@ -46,12 +46,14 @@ cat >"$stubdir/x-www-browser" <<EOF
echo "stub browser invoked with: \$1" >&2
# Also fetch an option page and require a rendered title='' tooltip: proves the
# option template expands and the \${html:} filter escapes into the attribute.
# option9 additionally proves the WARC control renders with its expanded label.
# option9/option8 additionally prove the WARC and sitemap controls render.
opturl="\${1%/}/server/option2.html"
warcurl="\${1%/}/server/option9.html"
smurl="\${1%/}/server/option8.html"
if body="\$(curl -fsSL --max-time 20 "\$1")" && printf '%s' "\$body" | grep -qai httrack && printf '%s' "\$body" | grep -qaF step2.html &&
opt="\$(curl -fsSL --max-time 20 "\$opturl")" && printf '%s' "\$opt" | grep -qaF "title='" &&
warc="\$(curl -fsSL --max-time 20 "\$warcurl")" && printf '%s' "\$warc" | grep -qaF 'name="warcfile"' && printf '%s' "\$warc" | grep -qaF WARC; then
warc="\$(curl -fsSL --max-time 20 "\$warcurl")" && printf '%s' "\$warc" | grep -qaF 'name="warcfile"' && printf '%s' "\$warc" | grep -qaF WARC &&
sm="\$(curl -fsSL --max-time 20 "\$smurl")" && printf '%s' "\$sm" | grep -qaF 'name="sitemapurl"' && printf '%s' "\$sm" | grep -qaF 'name="sitemap"'; then
echo PASS >"$marker"
else
echo "FAIL: unexpected response from \$1" >"$marker"