|
|
|
|
@@ -219,6 +219,23 @@ that a <tt>403 Forbidden</tt> is a server refusal, not a robots rule: robots
|
|
|
|
|
options will not help there. That is an
|
|
|
|
|
<a href="#identity">identity</a> problem.</p>
|
|
|
|
|
|
|
|
|
|
<h4>Filter wildcards</h4>
|
|
|
|
|
<p>Inside a filter pattern, <tt>*</tt> matches any run of characters; a few
|
|
|
|
|
bracket forms match narrower sets. The full table, with size and mime rules, is on
|
|
|
|
|
<a href="filters.html">the filters page</a>.</p>
|
|
|
|
|
<table class="tblRegular tableWidth" border="0">
|
|
|
|
|
<tr class="tblHeaderColor"><td><b>Wildcard</b></td><td><b>Matches</b></td><td><b>Example</b></td></tr>
|
|
|
|
|
<tr><td><tt>*</tt></td><td>any run of characters</td><td><tt>+*.pdf</tt> — any URL ending <tt>.pdf</tt></td></tr>
|
|
|
|
|
<tr><td><tt>*[file]</tt>, <tt>*[name]</tt></td><td>one path segment (any char but <tt>/</tt> and <tt>?</tt>)</td><td><tt>example.com/*[file]/</tt> — a directory-index page</td></tr>
|
|
|
|
|
<tr><td><tt>*[path]</tt></td><td>a path, slashes allowed (any char but <tt>?</tt>)</td><td><tt>example.com/*[path].zip</tt></td></tr>
|
|
|
|
|
<tr><td><tt>*[param]</tt></td><td>an optional query string</td><td><tt>page.html*[param]</tt> matches with or without <tt>?...</tt></td></tr>
|
|
|
|
|
<tr><td><tt>*[a,b,c]</tt></td><td>any one character in the set</td><td><tt>*[a,b,c].txt</tt></td></tr>
|
|
|
|
|
<tr><td><tt>*[a-z]</tt></td><td>any one character in the range</td><td><tt>img*[0-9].gif</tt></td></tr>
|
|
|
|
|
<tr><td><tt>*[\x]</tt></td><td>the literal character x (escapes <tt>* [ ] \</tt>)</td><td><tt>*[\*]</tt> matches a real <tt>*</tt></td></tr>
|
|
|
|
|
<tr><td><tt>*[<NN]</tt>, <tt>*[>NN]</tt></td><td>file size in KB below / above NN</td><td><tt>-*.gif*[<5]</tt> skips GIFs under 5 KB</td></tr>
|
|
|
|
|
<tr><td><tt>*[]</tt></td><td>end anchor: nothing may follow</td><td><tt>*.html*[]</tt> rejects <tt>i.html?p=1</tt></td></tr>
|
|
|
|
|
</table>
|
|
|
|
|
|
|
|
|
|
<h3 id="limits">4. Limits and politeness</h3>
|
|
|
|
|
|
|
|
|
|
<p>HTTrack ships cautious on purpose: it is easy to hammer a small site by
|
|
|
|
|
@@ -404,14 +421,16 @@ host and stops the crawl; start from the final URL, or add
|
|
|
|
|
rule wins.</small></p>
|
|
|
|
|
|
|
|
|
|
<h4>Download the PDFs on a site</h4>
|
|
|
|
|
<p><tt>httrack https://example.com/ --path mydir</tt><br>
|
|
|
|
|
<small>There is no PDF-only crawl. HTTrack discovers PDF links by parsing the
|
|
|
|
|
site's HTML pages, so a <tt>"-*" "+example.com/*.pdf"</tt> filter blocks the very
|
|
|
|
|
pages that carry the links and grabs only the PDFs linked from the front page. Let
|
|
|
|
|
it crawl the site: the HTML pages come along as the scaffolding, and every reachable
|
|
|
|
|
PDF is saved with them. If some PDFs live on another host (a CDN or a docs
|
|
|
|
|
subdomain), allow that host too, for example
|
|
|
|
|
<tt>"+docs.example.com/*.pdf"</tt>.</small></p>
|
|
|
|
|
<p><tt>httrack https://example.com/ "-*" "+https://example.com/*.html" "+https://example.com/*[path]/" "+https://example.com/*.pdf" --path mydir</tt><br>
|
|
|
|
|
<small>HTTrack finds PDFs by parsing the site's HTML, so a plain
|
|
|
|
|
<tt>"-*" "+example.com/*.pdf"</tt> is wrong: it prunes the pages that carry the
|
|
|
|
|
links and keeps only PDFs reachable from the front page. Instead admit the HTML as
|
|
|
|
|
scaffolding (<tt>*.html</tt> and <tt>*[path]/</tt> for directory-index pages at any
|
|
|
|
|
depth, e.g. <tt>docs/</tt> or <tt>a/b/deep/</tt>; <tt>*[file]/</tt> would stop at one
|
|
|
|
|
level), keep the PDFs, and let <tt>-*</tt> drop everything else (images,
|
|
|
|
|
archives, off-site assets). PDFs on another host (a CDN or docs subdomain) are not
|
|
|
|
|
included by default; allow that host too, e.g. <tt>"+docs.example.com/*.pdf"</tt>,
|
|
|
|
|
or widen to <tt>"+*.pdf"</tt> for PDFs anywhere.</small></p>
|
|
|
|
|
|
|
|
|
|
<h4>Keep page requisites, including off-host images</h4>
|
|
|
|
|
<p><tt>httrack https://example.com/blog/ --near --path mydir</tt><br>
|
|
|
|
|
|