Compare commits

...

7 Commits

Author SHA1 Message Date
Xavier Roche
99bd9bb674 A failed re-fetch overwrote the mirrored file with the aborted read's debris (#763)
* A failed re-fetch overwrote the mirrored file with the aborted read's debris

A transfer that dies before a complete response has no body, yet the save path
still consulted r.adr, which at that point holds whatever the aborted header
read left behind: raw status-line bytes, or an empty buffer that truncated the
file to zero on macOS. Require a successful transfer, as the empty-body half of
the condition already did.

Closes #748

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: assert the failed re-fetch keeps its bytes on every platform

Test 93 filtered reset.bin out of its bucket lists because a connection killed
before the status line surfaced differently per platform. It no longer does, so
assert the resource like any other: unchanged, with the bytes pass 1 mirrored.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: make the re-fetch assertions catch a crawl that never re-fetched

The new test passed with no second pass at all, so a regression that stopped
re-fetching would have looked green. Assert the failure the fixture provokes,
give stay.bin a fresh pass-2 body so a fix that stopped overwriting anything
fails, and compare reset.bin by checksum rather than by length.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: assert the retry exhaustion, not the per-platform failure message

The message a cut connection produces depends on whether any bytes arrived, so
matching it would fail on a runner that sees none.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 09:44:28 +02:00
Xavier Roche
29dfd2df59 htsserver never returns from main(), it blocks on its own exit wait (#757)
htsthread_wait_n(background_threads - 1) subtracts one more than the count of
threads that must not be joined. Without --ppid that count is zero, so the wait
asks for a negative number of outstanding threads and the counter never gets
there; with --ppid it waits on the pinger, which by design never returns.

Wait for background_threads instead, which is what the sibling call inside
back_launch_cmd() already passes.

Closes #753

Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 09:43:51 +02:00
Xavier Roche
55765815f5 htsAddLink walks back from an empty codebase, one byte before the buffer (#767)
* htsAddLink walks back from an empty codebase, one byte before the buffer

Same idiom as the lienrelatif() underflow fixed in #729: the trim that walks
back to the last '/' starts at codebase + strlen(codebase) - 1, which is
codebase - 1 when the string is empty, and the loop dereferences it before
a > codebase stops it.

Unlike #729 there is no reachable empty value. codebase is copied from a
recorded link's fil, and no hts_record_link() call site can supply an empty
one: every fil is either a literal seed or an ident_url_absolute() /
ident_url_relatif() success return, and fil_simplifie() restores "/" or "./"
rather than leaving a path empty. The guard goes in anyway, and -#test=addlink
drives the walk directly: under the sanitize job's ASan+UBSan build the test
fails on the unfixed walk and passes with the guard. Its second case pins the
ordinary trim, which the guard leaves alone.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* review: add the case that actually notices the codebase trim

For an ordinary relative link ident_url_relatif() re-derives the directory
from the path it is handed, so deleting the trim outright left the two
existing cases green. A query-only link ("?x=1") copies that path whole, and
does catch it.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 09:40:11 +02:00
Xavier Roche
869b8479e9 structcheck() builds its rename target with an unbounded sprintf (#762)
* structcheck() builds its rename target with an unbounded sprintf

Both structcheck() and structcheck_utf8() move a regular file sitting where a
directory belongs, and build the "<name>.txt" target with a raw sprintf into a
2048-byte buffer. It stays in bounds only because of a strlen(path) >
HTS_URLMAXSIZE guard dozens of lines above, which nothing at the write site
mentions. Route both through sprintfbuff() and fail with ENAMETOOLONG, so the
bound is local.

Armed the probe: with that distant guard patched out, a 2045-byte path makes
the old sprintf write 2049 bytes into tmpbuf[2048] under ASan; the same build
with sprintfbuff() returns -1 and reports nothing.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* review: tighten the structcheck self-test and its comments

The path builder could end a path with a bare separator when the base
directory's length hit the wrong residue, so the test aborted on a long
$TMPDIR. The refusal case also asserted on a component structcheck never
creates; assert on the outermost one instead.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* tests: drop the max-length rename case, macOS PATH_MAX is 1024

The path the guard admits (HTS_URLMAXSIZE) plus ".txt" is longer than macOS
accepts, so fopen() failed there. The case could not tell a fixed build from an
unfixed one anyway; what is left covers the guard and the rename on both entry
points.

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 09:34:02 +02:00
Xavier Roche
9fe47c3986 Remove dead HTS_USEZLIB guards now that zlib is mandatory (#761)
configure has rejected --without-zlib since #750, and htsglobal.h
#errors on a forced HTS_USEZLIB=0, so the macro can only ever be 1.
Collapse every #if/#ifdef HTS_USEZLIB guard (htsweb.c, htsname.c,
htslib.c, htszlib.c, htsselftest.c, htscodec.c) to its always-taken
branch; htsweb.c's guard used #ifdef where every other site used #if,
an inconsistency that no longer matters once the guard is gone.

htscodec.c's #else arms were a genuine zlib-free content-coding path
(Accept-Encoding: identity, hts_codec_unpack returning -1), not stubs.
Removed for consistency with the other nine guards: the cache
(htscache.c, proxy/store.c) already reaches minizip unconditionally
from ~60 call sites with no null backend, so a zlib-free build is not
actually reachable today regardless of this file.

No new test: this is dead-code removal with no behavior change.
Verified by differential build against master: htscodec.o and
htszlib.o are byte-identical; the other touched objects differ only
in __LINE__ immediates shifted by the removed guard lines (plus one
cosmetic objdump label-annotation artifact each in htsselftest.o and
htsweb.o, from string-literal pool reordering). Exported symbols in
libhttrack.so are unchanged. make check: 166/166 (158 pass, 8 skip).

Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 09:05:52 +02:00
Xavier Roche
37fa549ac5 Add the regression test #747 could not carry when it landed (#760)
htsselftest.c and tests/Makefile.am were held by #718 while #747 was fixed, so
the thread-counting fix went in without a test. -#test=threadwait covers it
from both sides: a wait placed right after a spawn joins that thread, and
wait_n(n) leaves n running rather than draining them.

One spawn per round is what makes it bite. A batch gives the earlier threads
time to raise the counter themselves, which is why an eight-thread version
passed on the unfixed engine; one thread per round failed 10 runs out of 10.

The changes-race self-test can now drop the counter it kept because
htsthread_wait() could not be trusted to join.

Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 09:05:44 +02:00
Xavier Roche
de8c0eebfc Three hand-rolled copies of the unlink-then-rename fallback, with two return conventions (#754)
* One unlink-then-rename helper for the three copies

replace_file() (htsback.c), wacz_rename_over() (htswarc.c) and the inline
block in singlefile_rewrite_file() each worked around Windows' non-clobbering
rename, with two opposite return conventions and only one of the three
converting path separators. Copying the wrong one gets you an inverted success
test.

hts_rename_over() replaces all four call sites: hts_boolean return, fconv() on
both paths. It lands in htsname.c rather than the htstools.c the issue named,
to stay clear of an in-flight PR over that file.

The wacz test now runs a cacheless second pass over a poisoned copy of the
package the first pass wrote, so repackaging has to clobber an existing file.
43 and 74 also assert the failure warnings stay out of the log, which is all
an inverted return leaves behind at those two sites.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <xroche@gmail.com>

* rename helper: move hts_rename_over to htstools.c

Issue #726 asked for htstools.c; it went to htsname.c only because #718 held
the file at the time. Kept internal, not HTSEXT_API: htscore.h pulls in both
headers, so every call site reaches it either way.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

* format: drop the trailing blank line left by the move

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Xavier Roche <roche@httrack.com>

---------

Signed-off-by: Xavier Roche <xroche@gmail.com>
Signed-off-by: Xavier Roche <roche@httrack.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 08:58:20 +02:00
23 changed files with 506 additions and 110 deletions

View File

@@ -585,14 +585,6 @@ static int create_back_tmpfile(httrackp *opt, lien_back *const back,
return 0;
}
/* Move src onto dst; RENAME does not clobber an existing target on Windows. */
static hts_boolean replace_file(const char *src, const char *dst) {
if (RENAME(src, dst) == 0)
return HTS_TRUE;
(void) UNLINK(dst);
return RENAME(src, dst) == 0 ? HTS_TRUE : HTS_FALSE;
}
/* Commit or restore a re-fetch backup (#77 follow-up): a re-fetch over an
existing file moved the good copy to back->tmpfile before truncating url_sav.
commit keeps the new file and drops the backup; else restore it so an aborted
@@ -610,7 +602,7 @@ static void back_finalize_backup(httrackp *opt, lien_back *const back,
}
/* On failure keep the backup: an orphaned temp beats losing the good copy.
*/
if (!replace_file(back->tmpfile, back->url_sav))
if (!hts_rename_over(back->tmpfile, back->url_sav))
hts_log_print(opt, LOG_WARNING | LOG_ERRNO,
"could not restore %s; previous copy kept as %s",
back->url_sav, back->tmpfile);
@@ -741,7 +733,7 @@ int back_finalize(httrackp * opt, cache_back * cache, struct_back * sback,
"Read error when decompressing");
}
UNLINK(unpacked);
} else if (replace_file(unpacked, back[p].url_sav)) {
} else if (hts_rename_over(unpacked, back[p].url_sav)) {
/* The temp bypassed filecreate(), which is what chmods. */
#ifndef _WIN32
chmod(back[p].url_sav, HTS_ACCESS_FILE);

View File

@@ -72,7 +72,7 @@ hts_codec hts_codec_parse(const char *encoding) {
return HTS_CODEC_IDENTITY;
if (strfield2(encoding, "gzip") || strfield2(encoding, "x-gzip") ||
strfield2(encoding, "deflate") || strfield2(encoding, "x-deflate"))
return HTS_USEZLIB ? HTS_CODEC_DEFLATE : HTS_CODEC_UNSUPPORTED;
return HTS_CODEC_DEFLATE;
if (strfield2(encoding, "br"))
return HTS_USEBROTLI ? HTS_CODEC_BROTLI : HTS_CODEC_UNSUPPORTED;
if (strfield2(encoding, "zstd"))
@@ -98,16 +98,11 @@ hts_codec hts_codec_parse(const char *encoding) {
const char *hts_acceptencoding(hts_boolean compressible, hts_boolean secure) {
if (!compressible)
return "identity";
#if HTS_USEZLIB
/* br and zstd over TLS only, as browsers do: a cleartext intermediary that
rewrites a coding it can not read would corrupt the mirror. */
if (secure)
return "gzip, deflate" HTS_AE_BROTLI HTS_AE_ZSTD ", identity;q=0.9";
return "gzip, deflate, identity;q=0.9";
#else
(void) secure;
return "identity";
#endif
}
hts_boolean hts_codec_is_archive_ext(hts_codec codec, const char *ext) {
@@ -300,11 +295,7 @@ int hts_codec_unpack(hts_codec codec, const char *filename,
return -1;
switch (codec) {
case HTS_CODEC_DEFLATE:
#if HTS_USEZLIB
return hts_zunpack(filename, newfile);
#else
return -1;
#endif
case HTS_CODEC_BROTLI:
case HTS_CODEC_ZSTD:
break;
@@ -348,10 +339,8 @@ size_t hts_codec_head(hts_codec codec, const void *in, size_t in_len, void *out,
if (in == NULL || in_len == 0 || out == NULL || out_len == 0)
return 0;
switch (codec) {
#if HTS_USEZLIB
case HTS_CODEC_DEFLATE:
return hts_zhead(in, in_len, out, out_len);
#endif
#if HTS_USEBROTLI
case HTS_CODEC_BROTLI:
return codec_head_brotli(in, in_len, out, out_len);

View File

@@ -1980,10 +1980,9 @@ int httpmirror(char *url1, httrackp * opt) {
}
// ATTENTION C'EST ICI QU'ON SAUVE LE FICHIER!!
// An empty body must not overwrite the file when the transfer failed
// (statuscode <= 0, e.g. an -M hard-stop): it would truncate a good
// copy to 0 (#77 follow-up).
if (r.adr != NULL || (r.size == 0 && r.statuscode > 0)) {
// A failed transfer has no body: r.adr holds debris from the aborted
// read, which would destroy the copy being re-fetched (#748).
if (r.statuscode > 0 && (r.adr != NULL || r.size == 0)) {
file_notify(opt, urladr(), urlfil(), savename(), 1, 1, r.notmodified);
if (filesave(opt, r.adr, (int) r.size, savename(), urladr(), urlfil()) !=
0) {
@@ -2693,7 +2692,11 @@ HTSEXT_API int structcheck(const char *path) {
if (!S_ISDIR(st.st_mode)) {
#if HTS_REMOVE_ANNOYING_INDEX
if (S_ISREG(st.st_mode)) { /* Regular file in place ; move it and create directory */
sprintf(tmpbuf, "%s.txt", file);
/* bounded here, not by the path-length guard far above */
if (!sprintfbuff(tmpbuf, "%s.txt", file)) {
errno = ENAMETOOLONG;
return -1;
}
if (rename(file, tmpbuf) != 0) { /* Can't rename regular file */
return -1;
}
@@ -2801,7 +2804,11 @@ HTSEXT_API int structcheck_utf8(const char *path) {
if (!S_ISDIR(st.st_mode)) {
#if HTS_REMOVE_ANNOYING_INDEX
if (S_ISREG(st.st_mode)) { /* Regular file in place ; move it and create directory */
sprintf(tmpbuf, "%s.txt", file);
/* bounded here, not by the path-length guard far above */
if (!sprintfbuff(tmpbuf, "%s.txt", file)) {
errno = ENAMETOOLONG;
return -1;
}
if (RENAME(file, tmpbuf) != 0) { /* Can't rename regular file */
return -1;
}
@@ -3821,11 +3828,12 @@ int htsAddLink(htsmoduleStruct * str, char *link) {
strcpybuff(codebase, heap(ptr)->fil);
else
strcpybuff(codebase, heap(heap(ptr)->precedent)->fil);
a = codebase + strlen(codebase) - 1;
// empty codebase has no last char; codebase-1 would underflow
a = codebase[0] != '\0' ? codebase + strlen(codebase) - 1 : codebase;
while((*a) && (*a != '/') && (a > codebase))
a--;
if (*a == '/')
*(a + 1) = '\0'; // couper
*(a + 1) = '\0'; // cut
} else { // couper http:// éventuel
if (strfield(codebase, "http://")) {
char BIGSTK tempo[HTS_URLMAXSIZE * 2];

View File

@@ -1131,12 +1131,10 @@ int http_sendhead(httrackp * opt, t_cookie * cookie, int mode,
// Compression accepted ?
if (retour->req.http11) {
hts_boolean compressible = HTS_FALSE;
hts_boolean compressible =
(!retour->req.range_used && !retour->req.nocompression);
hts_boolean secure = HTS_FALSE;
#if HTS_USEZLIB
compressible = (!retour->req.range_used && !retour->req.nocompression);
#endif
#if HTS_USEOPENSSL
secure = retour->ssl ? HTS_TRUE : HTS_FALSE;
#endif

View File

@@ -43,9 +43,7 @@ Please visit our Website: http://www.httrack.com
#include "htsencoding.h"
#include "htssniff.h"
#include "htscodec.h"
#if HTS_USEZLIB
#include "htszlib.h"
#endif
#include <ctype.h>
#include <limits.h>

View File

@@ -43,6 +43,7 @@ Please visit our Website: http://www.httrack.com
#include "htsglobal.h"
#include "htscore.h"
#include "htsmodules.h"
#include "htsback.h"
#include "htsdefines.h"
#include "htslib.h"
@@ -63,9 +64,7 @@ Please visit our Website: http://www.httrack.com
#include "htswarc.h"
#include "htschanges.h"
#include "htssinglefile.h"
#if HTS_USEZLIB
#include "htszlib.h"
#endif
#if HTS_USEZSTD
#include <zstd.h>
#endif
@@ -2623,6 +2622,68 @@ static int st_savename(httrackp *opt, int argc, char **argv) {
return 0;
}
/* an empty fil started htsAddLink's codebase walk before the buffer (#730) */
static int st_addlink(httrackp *opt, int argc, char **argv) {
htsmoduleStruct BIGSTK str;
cache_back cache;
struct_back *sback;
hash_struct hash;
int ptr = 0;
int i;
(void) argc;
(void) argv;
memset(&cache, 0, sizeof(cache));
cache.hashtable = (void *) coucal_new(0);
sback = back_new(opt, opt->maxsoc * 32 + 1024);
/* same wiring as hts_mirror (htscore.c) */
hash_init(opt, &hash, opt->urlhack);
hash.liens = (const lien_url *const *const *) &opt->liens;
opt->hash = &hash;
hts_record_init(opt);
memset(&str, 0, sizeof(str));
str.opt = opt;
str.sback = sback;
str.cache = &cache;
str.hashptr = &hash;
str.ptr_ = &ptr;
str.addLink = htsAddLink;
/* [0] is the underflow; [1] and [2] are controls that the trim is unchanged.
A query-only link is the one that notices the trim at all: for the others
ident_url_relatif() re-derives the directory from the path it is given. */
for (i = 0; i < 3; i++) {
static const char *const fil[3] = {"", "/dir/page.html", "/dir/page.html"};
static const char *const lnk[3] = {"sub/page.html", "sub/page.html",
"?x=1"};
static const char *const want[3] = {
"untouched", "http://www.example.com/dir/sub/page.html",
"http://www.example.com/dir/?x=1"};
char BIGSTK loc[HTS_URLMAXSIZE * 2];
char BIGSTK link[HTS_URLMAXSIZE];
strcpybuff(loc, "untouched");
strcpybuff(link, lnk[i]);
str.localLink = loc;
str.localLinkSize = (int) sizeof(loc);
if (!hts_record_link(opt, "www.example.com", fil[i], "", "", "", ""))
return 1;
ptr = heap_top_index();
str.url_host = heap(ptr)->adr;
str.url_file = heap(ptr)->fil;
assertf(htsAddLink(&str, link) == 0); /* refused by the wizard either way */
if (strcmp(loc, want[i]) != 0) {
fprintf(stderr, "addlink[%d]: got '%s' want '%s'\n", i, loc, want[i]);
return 1;
}
}
printf("addlink self-test OK\n");
return 0;
}
static int st_cache(httrackp *opt, int argc, char **argv) {
int err;
@@ -3264,6 +3325,81 @@ static int st_topindex(httrackp *opt, int argc, char **argv) {
return 0;
}
/* Build a path of exactly len chars under base; returns that length. */
static size_t st_structcheck_longpath(char *dst, size_t dstsize,
const char *base, size_t len) {
size_t n = strlen(base);
assertf(len < dstsize && n + 2 <= len);
memmove(dst, base, n);
while (n < len) {
size_t seg = len - n - 1;
if (seg > 200) /* stay under the usual 255-byte component limit */
seg = len - n == 202 ? 199 : 200; /* never leave a bare separator */
dst[n++] = '/';
memset(dst + n, 'x', seg);
n += seg;
}
dst[n] = '\0';
return n;
}
/* The path guard, and the <name>.txt rename structcheck() performs when a
regular file sits where a directory has to go (#745). */
static int st_structcheck(httrackp *opt, int argc, char **argv) {
char BIGSTK path[HTS_URLMAXSIZE * 2];
char BIGSTK target[HTS_URLMAXSIZE * 2];
FILE *fp;
(void) opt;
if (argc < 1) {
fprintf(stderr, "usage: -#test=structcheck <writable directory>\n");
return 1;
}
/* over the guard: refused before a single directory is created */
st_structcheck_longpath(path, sizeof(path), argv[0], HTS_URLMAXSIZE + 1);
errno = 0;
assertf(structcheck(path) == -1);
assertf(errno == EINVAL);
errno = 0;
assertf(structcheck_utf8(path) == -1);
assertf(errno == EINVAL);
{
char *const sep = strchr(path + strlen(argv[0]) + 1, '/');
assertf(sep != NULL);
sep[1] = '\0'; /* the outermost component it would have created */
assertf(!dir_exists(path));
}
/* a regular file where a directory belongs is renamed to <name>.txt */
snprintf(path, sizeof(path), "%s/sc", argv[0]);
fp = fopen(path, "wb");
assertf(fp != NULL);
fclose(fp);
snprintf(path, sizeof(path), "%s/sc/sub/", argv[0]);
assertf(structcheck(path) == 0);
assertf(dir_exists(path));
snprintf(target, sizeof(target), "%s/sc.txt", argv[0]);
assertf(fexist(target));
/* the utf-8 entry point carries the same rename */
snprintf(path, sizeof(path), "%s/u8", argv[0]);
fp = FOPEN(path, "wb");
assertf(fp != NULL);
fclose(fp);
snprintf(path, sizeof(path), "%s/u8/sub/", argv[0]);
assertf(structcheck_utf8(path) == 0);
assertf(dir_exists(path));
snprintf(target, sizeof(target), "%s/u8.txt", argv[0]);
assertf(fexist_utf8(target));
printf("structcheck self-test OK\n");
return 0;
}
/* Each inplace_escape_*() must equal escape_*() on a copy. */
static int st_inplace_escape(httrackp *opt, int argc, char **argv) {
/* >255 bytes forces the helper's malloct path, not the stack buffer */
@@ -3377,7 +3513,6 @@ static int st_status(httrackp *opt, int argc, char **argv) {
return 0;
}
#if HTS_USEZLIB
/* Deflate src->path at windowBits (16+ gzip, + zlib, - raw); 0 on success. */
static int ae_write_packed(const char *path, int windowBits,
const unsigned char *src, size_t len) {
@@ -3451,7 +3586,6 @@ static int ae_write_collision(const char *path, const unsigned char *src,
freet(buf);
return ok ? 0 : 1;
}
#endif
/* Write src[0..len) to path as-is; 0 on success. */
static int ae_write_raw(const char *path, const unsigned char *src,
@@ -3501,7 +3635,6 @@ static int st_acceptencoding(httrackp *opt, int argc, char **argv) {
assertf(strstr(on, "br") == NULL && strstr(on, "zstd") == NULL);
assertf((strstr(tls, ", br") != NULL) == (HTS_USEBROTLI != 0));
assertf((strstr(tls, "zstd") != NULL) == (HTS_USEZSTD != 0));
#if HTS_USEZLIB
if (argc >= 1) {
static const int windowBits[] = {16 + MAX_WBITS, MAX_WBITS, -MAX_WBITS};
const unsigned char small[] =
@@ -3580,10 +3713,6 @@ static int st_acceptencoding(httrackp *opt, int argc, char **argv) {
}
freet(body);
}
#else
(void) argc;
(void) argv;
#endif
printf("acceptencoding self-test OK: %s\n", on);
return 0;
}
@@ -3928,7 +4057,6 @@ static int st_sitemap(httrackp *opt, int argc, char **argv) {
freet(big);
}
#if HTS_USEZLIB
/* A highly compressible document decodes without running away: the ratio
budget cannot bind (deflate tops out near 1032:1), so this pins the
decompression path itself rather than the 64 MiB ceiling. */
@@ -3979,13 +4107,11 @@ static int st_sitemap(httrackp *opt, int argc, char **argv) {
freet(z);
freet(x);
}
#endif
/* An unterminated <loc> at end of buffer must not read past it. */
assertf(sm_scan("<urlset><loc>http://h.test/a", 100, &idx, &c) == 0);
assertf(sm_scan("<urlset><lo", 100, &idx, &c) == 0);
#if HTS_USEZLIB
/* A gzip-framed document is decompressed before scanning. */
{
const char *const xml =
@@ -4015,7 +4141,6 @@ static int st_sitemap(httrackp *opt, int argc, char **argv) {
assertf(hts_sitemap_scan(z, 4, 100, &idx, sm_take, &c) == -1);
freet(z);
}
#endif
/* robots.txt: only Sitemap: records, comments stripped, case-insensitive,
and group-independent (no User-agent line needed). */
@@ -6089,6 +6214,116 @@ static int st_changes(httrackp *opt, int argc, char **argv) {
return err;
}
/* #747: a thread is outstanding from the moment hts_newthread() returns, not
from the moment it starts running, or a wait right after the spawn joins
nothing. One thread per round is what makes the old bug visible: the wait
had to find the counter at zero, and a batch of spawns gives the earlier
threads time to raise it. Unfixed, one round in two caught it, so the round
count is what turns that into a reliable failure. */
#define THREADWAIT_N 8
#define THREADWAIT_ROUNDS 16
#define THREADWAIT_SLEEP_MS 50
#define THREADWAIT_GATE_MS 10000
static htsmutex threadwait_lock = HTSMUTEX_INIT;
static int threadwait_done = 0;
static hts_boolean threadwait_gated = HTS_FALSE;
static int threadwait_count(void) {
int n;
hts_mutexlock(&threadwait_lock);
n = threadwait_done;
hts_mutexrelease(&threadwait_lock);
return n;
}
static void threadwait_thread(void *arg) {
(void) arg;
Sleep(THREADWAIT_SLEEP_MS);
hts_mutexlock(&threadwait_lock);
threadwait_done++;
hts_mutexrelease(&threadwait_lock);
}
/* Stays outstanding until the gate clears, so wait_n() can be asked to leave a
known number of live threads behind. Bounded: a wait_n() that wrongly drains
them would otherwise never return, and hang the suite instead of failing. */
static void threadwait_gated_thread(void *arg) {
int waited;
(void) arg;
for (waited = 0; waited < THREADWAIT_GATE_MS; waited += 10) {
hts_boolean gated;
hts_mutexlock(&threadwait_lock);
gated = threadwait_gated;
hts_mutexrelease(&threadwait_lock);
if (!gated)
break;
Sleep(10);
}
hts_mutexlock(&threadwait_lock);
threadwait_done++;
hts_mutexrelease(&threadwait_lock);
}
static int st_threadwait(httrackp *opt, int argc, char **argv) {
int err = 0;
int i, round;
(void) opt;
(void) argc;
(void) argv;
/* htsthread_wait() joins a thread spawned just before it */
for (round = 0; round < THREADWAIT_ROUNDS && !err; round++) {
hts_mutexlock(&threadwait_lock);
threadwait_done = 0;
hts_mutexrelease(&threadwait_lock);
if (hts_newthread(threadwait_thread, NULL) != 0) {
fprintf(stderr, "threadwait: cannot spawn\n");
return 1;
}
htsthread_wait();
if (threadwait_count() != 1) {
fprintf(stderr, "threadwait: round %d returned before the thread ran\n",
round);
err = 1;
}
}
/* htsthread_wait_n(n) leaves n behind rather than draining everything */
hts_mutexlock(&threadwait_lock);
threadwait_done = 0;
threadwait_gated = HTS_TRUE;
hts_mutexrelease(&threadwait_lock);
for (i = 0; i < THREADWAIT_N; i++) {
if (hts_newthread(threadwait_gated_thread, NULL) != 0) {
fprintf(stderr, "threadwait: cannot spawn a gated thread\n");
return 1;
}
}
htsthread_wait_n(THREADWAIT_N);
if (threadwait_count() != 0) {
fprintf(stderr, "threadwait: wait_n(%d) joined %d gated threads\n",
THREADWAIT_N, threadwait_count());
err = 1;
}
hts_mutexlock(&threadwait_lock);
threadwait_gated = HTS_FALSE;
hts_mutexrelease(&threadwait_lock);
htsthread_wait();
if (threadwait_count() != THREADWAIT_N) {
fprintf(stderr, "threadwait: wait left %d/%d gated threads running\n",
THREADWAIT_N - threadwait_count(), THREADWAIT_N);
err = 1;
}
printf("threadwait self-test: %s\n", err ? "FAIL" : "OK");
return err;
}
#define CHANGES_RACE_FILES 8
#define CHANGES_RACE_ROUNDS 400
@@ -6103,11 +6338,8 @@ static void changes_race_notify(httrackp *opt, int n) {
hts_changes_notify(opt, "race.example", fil, save, HTS_TRUE, HTS_FALSE);
}
/* htsthread_wait() counts a thread only once it is running, so it can return
before any of them started; join on our own counter instead. */
static htsmutex changes_race_lock = HTSMUTEX_INIT;
static int changes_race_started = 0;
static int changes_race_live = 0;
static int changes_race_count(int *which) {
int n;
@@ -6127,9 +6359,6 @@ static void changes_race_thread(void *arg) {
hts_mutexrelease(&changes_race_lock);
for (i = 0; i < CHANGES_RACE_ROUNDS; i++)
changes_race_notify(opt, i % CHANGES_RACE_FILES);
hts_mutexlock(&changes_race_lock);
changes_race_live--;
hts_mutexrelease(&changes_race_lock);
}
/* A transfer thread the crawl never joins (FTP) reaches hts_changes_notify()
@@ -6184,7 +6413,6 @@ static int st_changes_race(httrackp *opt, int argc, char **argv) {
threads reaching a fresh one together race on the init itself. */
hts_mutexlock(&changes_race_lock);
changes_race_started = 0;
changes_race_live = 4;
hts_mutexrelease(&changes_race_lock);
for (i = 0; i < 4; i++) {
if (hts_newthread(changes_race_thread, opt) != 0) {
@@ -6198,8 +6426,7 @@ static int st_changes_race(httrackp *opt, int argc, char **argv) {
for (i = 0; i < 64; i++)
hts_changes_report(opt, &out);
hts_changes_close_opt(opt);
while (changes_race_count(&changes_race_live) > 0)
Sleep(10);
htsthread_wait();
/* Sealed: a straggler must be dropped, not start a report nobody writes. */
changes_race_notify(opt, CHANGES_RACE_FILES + 1);
@@ -6277,6 +6504,8 @@ static const struct selftest_entry {
st_changes},
{"changes-race", "<dir>", "--changes under a late transfer thread (#714)",
st_changes_race},
{"threadwait", "", "htsthread_wait() joins threads spawned just before it",
st_threadwait},
{"pause", "", "randomized inter-file pause target self-test", st_pause},
{"relative", "<link> <curr-file>", "relative link between two paths",
st_relative},
@@ -6307,6 +6536,8 @@ static const struct selftest_entry {
{"fsize", "<dir>", "file size past the 2GB signed-32-bit wrap", st_fsize},
{"growsize", "", "buffer capacity for a 64-bit file size (no int wrap)",
st_growsize},
{"addlink", "", "htsAddLink codebase walk over an empty current path",
st_addlink},
{"cache", "<dir>", "cache read/write round-trip self-test", st_cache},
{"cacheindex", "", "cache-index (.ndx) parse must stay in bounds",
st_cacheindex},
@@ -6330,6 +6561,9 @@ static const struct selftest_entry {
{"useragent", "", "default User-Agent self-test", st_useragent},
{"makeindex", "[dir]", "hts_finish_makeindex footer/refresh self-test",
st_makeindex},
{"structcheck", "<dir>",
"structcheck path guard and the <name>.txt rename it performs",
st_structcheck},
{"topindex", "[dir]",
"hts_buildtopindex charset handling of a non-ASCII project dir",
st_topindex},

View File

@@ -1108,13 +1108,8 @@ hts_boolean singlefile_rewrite_file(httrackp *opt, const char *root,
(void) chmod(fconv(catbuff, sizeof(catbuff), StringBuff(tmp)),
HTS_ACCESS_FILE);
#endif
if (ok) {
/* RENAME does not clobber an existing target on Windows. */
if (RENAME(StringBuff(tmp), page_path) != 0) {
(void) UNLINK(page_path);
ok = RENAME(StringBuff(tmp), page_path) == 0 ? HTS_TRUE : HTS_FALSE;
}
}
if (ok)
ok = hts_rename_over(StringBuff(tmp), page_path);
if (!ok) {
hts_log_print(opt, LOG_ERROR, "single-file: could not rewrite %s",
page_path);

View File

@@ -1443,3 +1443,15 @@ HTSEXT_API hts_boolean hts_findissystem(find_handle find) {
}
return 0;
}
hts_boolean hts_rename_over(const char *src, const char *dst) {
char csrc[CATBUFF_SIZE], cdst[CATBUFF_SIZE];
fconv(csrc, sizeof(csrc), src);
fconv(cdst, sizeof(cdst), dst);
if (RENAME(csrc, cdst) == 0)
return HTS_TRUE;
/* RENAME does not clobber an existing target on Windows. */
(void) UNLINK(cdst);
return RENAME(csrc, cdst) == 0 ? HTS_TRUE : HTS_FALSE;
}

View File

@@ -137,6 +137,10 @@ HTSEXT_API hts_boolean hts_findisdir(find_handle find);
HTSEXT_API hts_boolean hts_findisfile(find_handle find);
HTSEXT_API hts_boolean hts_findissystem(find_handle find);
/* Move src onto dst, replacing an existing dst; HTS_TRUE on success. Both
paths are fconv()'d. */
hts_boolean hts_rename_over(const char *src, const char *dst);
#endif
#endif

View File

@@ -964,16 +964,6 @@ static zipFile wacz_zip_open(const char *path) {
return zipOpen2_64(path, 0 /*create*/, NULL, &ff);
}
/* Move src onto dst (UTF-8/Windows-safe); RENAME won't clobber on Windows, so
fall back to unlink+rename. Returns 0 on success. */
static int wacz_rename_over(const char *src, const char *dst) {
char cs[CATBUFF_SIZE], cd[CATBUFF_SIZE];
if (RENAME(fconv(cs, sizeof(cs), src), fconv(cd, sizeof(cd), dst)) == 0)
return 0;
(void) UNLINK(fconv(cd, sizeof(cd), dst));
return RENAME(fconv(cs, sizeof(cs), src), fconv(cd, sizeof(cd), dst));
}
/* Package the segment(s) + .cdx + a generated pages.jsonl into <base>.wacz at
crawl end (the archive file(s) and .cdx are already closed on disk). */
static void warc_wacz_package(warc_writer *w) {
@@ -1092,7 +1082,7 @@ static void warc_wacz_package(warc_writer *w) {
hts_log_print(w->opt, LOG_WARNING,
"WACZ: packaging failed, kept existing %s untouched",
waczpath);
} else if (wacz_rename_over(tmppath, waczpath) != 0) {
} else if (!hts_rename_over(tmppath, waczpath)) {
(void) UNLINK(fconv(catbuff, sizeof(catbuff), tmppath));
hts_log_print(w->opt, LOG_WARNING | LOG_ERRNO,
"WACZ: could not finalize %s", waczpath);

View File

@@ -101,8 +101,8 @@ static void htsweb_sig_brpipe(int code) {
/* ignore */
}
/* Number of background threads */
static int background_threads = 0;
/* Threads that never return; no wait may count on them draining. */
static int nonjoinable_threads = 0;
/* Server/client ping handling */
static htsmutex pingMutex = HTSMUTEX_INIT;
@@ -224,9 +224,7 @@ int main(int argc, char *argv[]) {
#ifdef HTS_USESWF
smallserver_setkey("USESWF", "1");
#endif
#ifdef HTS_USEZLIB
smallserver_setkey("USEZLIB", "1");
#endif
#ifdef _WIN32
smallserver_setkey("WIN32", "1");
#endif
@@ -299,15 +297,19 @@ int main(int argc, char *argv[]) {
/* pinger */
if (parentPid > 0) {
hts_newthread(client_ping, (void *) (uintptr_t) parentPid);
background_threads++; /* Do not wait for this thread! */
if (hts_newthread(client_ping, (void *) (uintptr_t) parentPid) == 0) {
#ifndef _WIN32
nonjoinable_threads++; /* client_ping() only ever leaves through exit() */
#endif
}
smallserver_setpinghandler(pingHandler, NULL);
}
/* launch */
ret = help_server(argv[1], defaultPort, bindAddr);
htsthread_wait_n(background_threads - 1);
/* Drain everything a mirror may still have in flight, the pinger aside. */
htsthread_wait_n(nonjoinable_threads);
hts_uninit();
#ifdef _WIN32
@@ -382,7 +384,6 @@ void webhttrack_main(char *cmd) {
commandRunning = 1;
DEBUG(fprintf(stderr, "commandRunning=1\n"));
hts_newthread(back_launch_cmd, (void *) strdup(cmd));
background_threads++; /* Do not wait for this thread! */
}
void webhttrack_lock(void) {
@@ -423,8 +424,8 @@ static int webhttrack_runmain(httrackp * opt, int argc, char **argv) {
/* Rock'in! */
ret = hts_main2(argc, argv, opt);
/* Wait for pending threads to finish */
htsthread_wait_n(background_threads);
/* Wait for pending threads to finish; the pinger and this thread stay. */
htsthread_wait_n(nonjoinable_threads + 1);
return ret;
}

View File

@@ -40,7 +40,6 @@ Please visit our Website: http://www.httrack.com
#include "htscodec.h"
#include "htszlib.h"
#if HTS_USEZLIB
/* zlib */
/*
#include <zlib.h>
@@ -274,4 +273,3 @@ const char *hts_get_zerror(int err) {
break;
}
}
#endif

View File

@@ -0,0 +1,7 @@
#!/bin/bash
#
# an empty current path underflowed htsAddLink's codebase walk (#730).
set -euo pipefail
httrack -O /dev/null -#test=addlink | grep -q "addlink self-test OK"

View File

@@ -0,0 +1,12 @@
#!/bin/bash
#
# structcheck(): the path-length guard, and the <name>.txt rename that once
# built its target with an unbounded sprintf (#745).
set -euo pipefail
dir=$(mktemp -d)
trap 'rm -rf "$dir"' EXIT
httrack -O /dev/null -#test=structcheck "$dir" |
grep -q "structcheck self-test OK"

View File

@@ -0,0 +1,28 @@
#!/bin/bash
#
set -euo pipefail
tmpdir=$(mktemp -d "${TMPDIR:-/tmp}/httrack_threadwait_st.XXXXXX") || exit 1
trap 'rm -rf "$tmpdir"' EXIT HUP INT QUIT PIPE TERM
# No pipe into grep: SIGPIPE would mask a failing exit status.
expect_ok() {
local label="$1" out
shift
out=$("$@" 2>&1) || {
echo "FAIL: ${label} exited non-zero: ${out}"
exit 1
}
case "$out" in
*"${label}: OK"*) ;;
*)
echo "FAIL: ${out}"
exit 1
;;
esac
}
# A thread is outstanding from the moment hts_newthread() returns, so a wait
# that follows the spawn joins it, and wait_n(n) still leaves n behind (#747).
expect_ok "threadwait self-test" httrack -O "${tmpdir}/o1" -#test=threadwait

View File

@@ -18,6 +18,7 @@ set -euo pipefail
bash "$top_srcdir/tests/local-crawl.sh" \
--rerun-args '--update -M400000' \
--log-found 'More than 400000 bytes have been transferred.. giving up' \
--log-not-found 'could not restore' \
--found bigtrunc/slow.bin \
--file-min-bytes bigtrunc/slow.bin 655360 \
--file-min-bytes bigtrunc/fast.bin 655360 \

View File

@@ -79,6 +79,29 @@ done
# 65535 is a valid port the old "< 65535" bound refused
web_accepted 65535
# --- htsserver returns from main() instead of blocking at exit -------------
# The exit wait must exclude exactly the threads that never return (#753).
# $tmp has no lang.def, so the server fails right after announcing and main()
# reaches the wait on its own; rc, not EXITED, carries the verdict.
web_exits() {
local rc=0
: >"$tmp/exit.log"
run_with_timeout 30 htsserver "$tmp" "$@" >"$tmp/exit.log" 2>&1 || rc=$?
grep -q "^EXITED" "$tmp/exit.log" ||
! echo "FAIL: #753: htsserver ${*:-(no options)} never finished serving" || exit 1
# exactly 1, not merely "not 124": a crash or an assertf abort also escapes
# the wait, and would otherwise read as a pass
test "$rc" -eq 1 ||
! echo "FAIL: #753: htsserver ${*:-(no options)} exited $rc, want 1" || exit 1
}
# no pinger: the excluded count went negative
web_exits
# a pinger, which never returns and so must stay excluded
web_exits --ppid $$
# --- proxytrack <proxy-addr:port> <ICP-addr:port> --------------------------
# A bad argument falls through to the usage screen; it had no range check at
# all, so 65616 quietly listened on port 80. A valid one binds and blocks.

View File

@@ -5,11 +5,16 @@
# real gate: STORE-mode entries, the fixed layout, recomputed sha256 digests and
# the datapackage-digest chain (plus py-wacz/pywb when importable). The crawl
# skips cleanly on a build without OpenSSL (no conformant SHA-256 -> no package).
#
# The cacheless second pass repackages over the .wacz the first one left: the
# clobber, not the happy rename, is what breaks when a platform refuses to
# rename onto an existing file (#726).
set -eu
: "${top_srcdir:=..}"
bash "$top_srcdir/tests/local-crawl.sh" --errors 0 --wacz-validate \
--rerun-args '-C0' --log-not-found 'could not finalize' \
--found 'mini304/index.html' --found 'mini304/page.html' \
httrack 'BASEURL/mini304/index.html' --warc-file warc-out --wacz

View File

@@ -94,10 +94,7 @@ expect_counts_match() {
echo "OK"
}
lines() { printf '%s\n' "$@" | sort; }
# A connection killed before the status line surfaces differently per platform
# (macOS truncates the file, #748; Linux leaves it), so reset.bin is asserted on
# its own and kept out of the exact lists.
listed_but_reset() { listed "$1" | grep -v '/reset\.bin$' | sort; }
digest() { cksum <"$1" | tr -d '[:space:]'; }
# --- pass 1: nothing to compare against --------------------------------------
httrack "${common[@]}" -O "$out" --purge-old=0 "${base}/changes/index.html" \
@@ -120,6 +117,7 @@ grep -aq "first crawl" "${out}/hts-log.txt" || {
}
host="127.0.0.1_${port}"
resetdigest=$(digest "${out}/${host}/changes/reset.bin")
# A leftover from some earlier failed attempt, at a name pass 2 fetches for the
# first time: on disk, but never part of the previous mirror, so it is new.
printf 'leftover junk' >"${out}/${host}/changes/fresh.html"
@@ -144,21 +142,19 @@ expect "gone is doomed.html and the old redirect target" \
expect "changed is exactly the four moved resources" \
"$(lines "${host}/changes/coded.bin" "${host}/changes/index.html" \
"${host}/changes/moved.bin" "${host}/changes/moved.html")" \
"$(listed_but_reset changed)"
"$(listed changed | sort)"
# Re-served byte for byte, 200, no Last-Modified, no ETag. flaky.bin gets there
# through a failed transfer and a retry, so its file is notified twice;
# redirtarget.html arrives behind a 302. sized.html has a byte-identical
# payload, but the rewritten link to the renamed redirect target makes the file
# on disk change length.
expect "unchanged is exactly the six stable resources" \
# on disk change length. reset.bin never completes a transfer (#746), so its
# previous copy stands.
expect "unchanged is exactly the seven stable resources" \
"$(lines "${host}/changes/codedstable.bin" "${host}/changes/flaky.bin" \
"${host}/changes/redirtarget.html" "${host}/changes/sized.html" \
"${host}/changes/stable.bin" "${host}/changes/stable.html")" \
"$(listed_but_reset unchanged)"
# reset.bin never completes a transfer, so it drops out of new.lst with its
# mirrored copy still there. Whatever else it is, it is not a deletion.
expect "a failed re-fetch is never reported gone" "" \
"$(listed gone | grep '/reset\.bin$' || true)"
"${host}/changes/redirtarget.html" "${host}/changes/reset.bin" \
"${host}/changes/sized.html" "${host}/changes/stable.bin" \
"${host}/changes/stable.html")" \
"$(listed unchanged | sort)"
expect_counts_match "pass 2 counts match its lists"
# Every mirrored file appears once and only once, across all four buckets.
@@ -192,19 +188,15 @@ test -f "${out}/${host}/changes/doomed.html" || {
exit 1
}
echo "OK"
printf '[a failed transfer keeps its file] ..\t'
test -f "${out}/${host}/changes/reset.bin" || {
echo "FAIL: reset.bin lost its previous copy"
exit 1
}
echo "OK"
expect "a failed transfer keeps its bytes" "$resetdigest" \
"$(digest "${out}/${host}/changes/reset.bin")"
# --- pass 3: the same run with purging on ------------------------------------
httrack "${common[@]}" -O "$out" --update "${base}/changes/index.html" \
>"${tmpdir}/log3" 2>&1
expect "the report says files were purged" "true" "$(field purged)"
expect "gone is the page that only pass 2 linked" \
"${host}/changes/transient.html" "$(listed_but_reset gone)"
"${host}/changes/transient.html" "$(listed gone | sort)"
printf '[a purged file is off disk] ..\t'
test ! -f "${out}/${host}/changes/transient.html" || {
echo "FAIL: transient.html survived --purge-old"

View File

@@ -0,0 +1,28 @@
#!/bin/bash
#
# A re-fetch cut mid-header must leave the mirrored file alone (#748): unfixed,
# the aborted read's leftover buffer lands on it. err.bin is the control, an
# HTTP 500 on the same resource already keeping its copy; stay.bin answers
# normally with a new body, so a fix that stopped overwriting anything fails.
#
# Purging is off: an update purge deletes both files for an unrelated reason
# (#746), which would hide what this asserts.
set -euo pipefail
: "${top_srcdir:=..}"
bash "$top_srcdir/tests/local-crawl.sh" \
--rerun-args '--update --purge-old=0' \
--log-found 'after 2 retries at link .*keep/page\.html' \
--log-found 'after 2 retries at link .*keep/data\.bin' \
--file-matches 'keep/page.html' 'KEEP-PAGE-V1' \
--file-not-matches 'keep/page.html' 'HTTP/1\.0 200' \
--file-matches 'keep/data.bin' 'KEEP-BIN-V1' \
--file-not-matches 'keep/data.bin' 'HTTP/1\.0 200' \
--file-min-bytes 'keep/data.bin' 2048 \
--file-matches 'keep/err.bin' 'KEEP-ERR-V1' \
--file-min-bytes 'keep/err.bin' 2048 \
--file-matches 'keep/stay.bin' 'KEEP-STAY-V2' \
--file-min-bytes 'keep/stay.bin' 2048 \
httrack 'BASEURL/keep/index.html' --retries=2

View File

@@ -30,6 +30,7 @@ TEST_EXTENSIONS = .test
TEST_LOG_COMPILER = $(BASH)
TESTS = \
00_runnable.test \
01_engine-addlink.test \
01_engine-changes.test \
01_engine-charset.test \
01_engine-cmdline.test \
@@ -81,7 +82,9 @@ TESTS = \
01_engine-status.test \
01_engine-stripquery.test \
01_engine-strsafe.test \
01_engine-structcheck.test \
01_engine-syscharset.test \
01_engine-threadwait.test \
01_engine-topindex.test \
01_engine-urlhack.test \
01_engine-unescape-bounds.test \
@@ -194,6 +197,7 @@ TESTS = \
92_local-proxytrack-ndx-fields.test \
93_local-changes.test \
94_local-single-file.test \
95_local-sitemap.test
95_local-sitemap.test \
96_local-refetch-keep.test
CLEANFILES = check-network_sh.cache

View File

@@ -57,6 +57,8 @@ rerun_dead=
tmpdir=
serverpid=
crawlpid=
wacz_poisoned=
wacz_poison="stale-wacz-that-a-second-pass-must-replace"
function warning {
echo "** $*" >&2
@@ -272,6 +274,14 @@ if test -n "$warc_validate"; then
test -z "$w1" || cp "$w1" "${tmpdir}/warc-pass1.gz"
fi
# Poison the first-pass .wacz when a second pass follows: repackaging moves the
# new archive over it, so the marker must be gone afterwards (#726). Poisoning
# beats comparing the two packages, which can come out byte-identical.
if test -n "$wacz_validate" && test -n "${rerun}${rerun_args}"; then
wacz_poisoned=$(find "$mirrorroot" -maxdepth 2 -name '*.wacz' 2>/dev/null | sort | tail -n1)
test -z "$wacz_poisoned" || echo "$wacz_poison" >"$wacz_poisoned"
fi
# --- optional second pass: re-mirror into the same dir (cache/update path) ----
if test -n "$rerun"; then
info "re-running httrack (update pass)"
@@ -416,6 +426,12 @@ if test -n "$wacz_validate"; then
fi
die "no .wacz file produced under $mirrorroot"
fi
if test -n "$wacz_poisoned"; then
info "checking the second pass replaced the .wacz"
grep -q "$wacz_poison" "$wacz_poisoned" 2>/dev/null &&
die "stale .wacz kept: $wacz_poisoned"
result "OK"
fi
validator=$(nativepath "${testdir}/wacz-validate.py")
info "validating WACZ package"
"$python" "$validator" "$(nativepath "$wacz")" >&2 || die "WACZ validation failed"

View File

@@ -1000,6 +1000,62 @@ class Handler(SimpleHTTPRequestHandler):
v = 1 if self.refetch_pass() == 1 else 2
self.send_raw(b"<html><body><p>STAY-V%d</p></body></html>" % v, "text/html")
# --- re-fetch cut mid-header, so nothing is stored (#746, #748) ---------
KEEP_PAGE = b"<html><body><p>KEEP-PAGE-V1</p></body></html>"
KEEP_BIN = b"KEEP-BIN-V1\n" + b"\x51\x52\x53\x54" * 512
KEEP_ERR = b"KEEP-ERR-V1\n" + b"\x61\x62\x63\x64" * 512
def send_cut_headers(self):
"""Hang up mid-header: the response never becomes parseable."""
self.close_connection = True
try:
self.wfile.write(b"HTTP/1.0 200 OK\r\nContent-Ty")
self.wfile.flush()
except OSError:
pass
self.connection.close()
def route_keep_index(self):
self.refetch_pass()
self.send_html(
'\t<a href="page.html">page</a>\n'
'\t<a href="data.bin">data</a>\n'
'\t<a href="err.bin">err</a>\n'
'\t<a href="stay.bin">stay</a>\n'
)
def route_keep_page(self):
if self.refetch_pass() == 1:
self.send_raw(self.KEEP_PAGE, "text/html")
else:
self.send_cut_headers()
def route_keep_data(self):
if self.refetch_pass() == 1:
self.send_raw(self.KEEP_BIN, "application/octet-stream")
else:
self.send_cut_headers()
# Control: an HTTP error on the same resource already keeps the copy, so
# the cut-header routes above must end up indistinguishable from it.
def route_keep_err(self):
if self.refetch_pass() == 1:
self.send_raw(self.KEEP_ERR, "application/octet-stream")
else:
self.send_response(500)
self.send_header("Content-Type", "text/html")
self.send_header("Content-Length", "0")
self.end_headers()
# Control: answers normally, with a new body on pass 2, so a fix that
# stopped overwriting mirrored files altogether would be caught.
def route_keep_stay(self):
v = 1 if self.refetch_pass() == 1 else 2
self.send_raw(
b"KEEP-STAY-V%d\n" % v + b"\x71\x72\x73\x74" * 512,
"application/octet-stream",
)
# Echo what httrack advertised, so a crawl can assert the header.
def route_codec_ae(self):
self.send_raw(
@@ -1950,6 +2006,11 @@ class Handler(SimpleHTTPRequestHandler):
"/uptrunc/page.html": route_uptrunc_page,
"/uptrunc/file.bin": route_uptrunc_file,
"/uptrunc/stay.html": route_uptrunc_stay,
"/keep/index.html": route_keep_index,
"/keep/page.html": route_keep_page,
"/keep/data.bin": route_keep_data,
"/keep/err.bin": route_keep_err,
"/keep/stay.bin": route_keep_stay,
"/types/index.html": route_types_index,
"/types/control.php": route_types,
"/types/photo.png": route_types,