785d8deadd
Wrote a new test to examine the redirect behavior of the crawler, ensuring that the redirect URL is the URL that is reported in the parquet file. This works as intended. Noticed in the course of this that the crawler doesn't add links from meta-tag redirects to the crawl frontier. Added logic to handle this case, amended the test case to verify the new behavior. Added the meta-redirect case to the HtmlDocumentProcessorPlugin as well, so that we consider it a link between documents in the unlikely case that a meta redirect is to another domain. |
||
---|---|---|
.. | ||
src/main/java/nu/marginalia/link_parser | ||
build.gradle | ||
readme.md |
Link Parser
Deals with the various cases in link parsing, such as relative links, internal links, external links, pathological links, etc.