CatgirlIntelligenceAgency

Author	SHA1	Message	Date
Viktor Lofgren	7b40c0bbee	(assistant) Clean up similar websites' results	2023-12-22 14:07:01 +01:00
Viktor Lofgren	dc773c5c20	(adjacencies) Clean up AdjacenciesLoader Make JDBC batching more consistent, also adds a test case for the loader.	2023-12-21 14:14:22 +01:00
Viktor Lofgren	b6253b03c2	(adjacencies) Fix bug in AdjacenciesLoader This fixes a bug where a prepared statement was created before the table it was supposed to insert into was created. This fails and does nothing. Furthermore, added the logging that would have warned about this failure, had it been in place.	2023-12-21 13:12:31 +01:00
Viktor Lofgren	a5bc29245b	(cleanup) Remove vestigial support for WARC crawl data streams	2023-12-20 15:46:21 +01:00
Viktor Lofgren	bfae478251	Refactor CrawlerRevisitor for better consistency	2023-12-20 15:21:49 +01:00
Viktor Lofgren	a7cd490593	(minor) Remove dead code.	2023-12-19 18:58:33 +01:00
Viktor Lofgren	283d2caa81	Merge remote-tracking branch 'origin/master'	2023-12-19 18:38:01 +01:00
Viktor Lofgren	dd8fb04886	(converter) Add sizeloadSizeAdvice field to several ProcessedDomain Since the sideloaders don't populate the documents list in ProcessedDomain to keep the memory footprint manageable, the code that estimates knownUrls etc. will set them to zero, which has negative effects on their ranking. This change will populate them with a bullshit value within a sane ballpark, ensuring that these domains show up in the rankings.	2023-12-19 18:37:51 +01:00
Viktor	ce8dca7659	Update Additional Contributors.md	2023-12-19 12:22:01 +01:00
Viktor	5bd3934d22	Merge pull request #64 from dreimolo/macos_AS_fix Macos apple silicon fix, and slight improvements to sample downloader	2023-12-18 18:29:14 +01:00
Viktor Lofgren	128f550ee5	(run) Download to a temporary file to avoid corruption from aborted downloads	2023-12-18 18:28:17 +01:00
Viktor Lofgren	3a56a06c4f	(warc) Add a fields for etags and last-modified headers to the new crawl data formats Make some temporary modifications to the CrawledDocument model to support both a "big string" style headers field like in the old formats, and explicit fields as in the new formats. This is a bit awkward to deal with, but it's a necessity until we migrate off the old formats entirely. The commit also adds a few tests to this logic.	2023-12-18 17:45:54 +01:00
Viktor Lofgren	126ac3816f	(converter) Reduce queue size in ConverterWriter The size of the ArrayBlockingQueue in ConverterWriter.java has been reduced from 4 to 1. This change aims to reduce the memory utilization by not having fully processed domains piling up in RAM. This may cause the writer to go idle in waiting for new data, but that may be preferable to an OOM.	2023-12-18 13:42:40 +01:00
Viktor Lofgren	d02bed1a55	(loader) Optimize DomainLoaderService for faster startups Initialization parameters in DomainLoaderService and DomainIdRegistry have been updated to improve performance. This is done by adding sane default sizes to the hash tables involved, reducing GC churn, but also by setting a sensible fetch size to the queries used, and not fetching irrelevant information such as the domain name.	2023-12-18 13:15:10 +01:00
Viktor Lofgren	b7ed0ce537	(loader) Reset count after executing batch in DomainLoaderService This should greatly speed up starting the loader process.	2023-12-18 12:43:53 +01:00
Viktor Lofgren	a742503508	(search) Add view for showing mutual links between two websites	2023-12-17 17:50:44 +01:00
Viktor Lofgren	33312ab09e	(geo-ip) Update readme	2023-12-17 16:08:33 +01:00
Viktor Lofgren	c422f0b9fb	(geo-ip) Tidy up error handling	2023-12-17 16:06:51 +01:00
Viktor	7797de80e3	Merge pull request #65 from MarginaliaSearch/asn-info Replace the ip2location-LITE IP geolocation data with ASN information from apnic.net	2023-12-17 15:04:29 +01:00
Viktor Lofgren	c92f1b8df8	(geo-ip) Revert removal of ip2location logic We do both ip2location and ASN data. The change also adds some keywords based on autonomous system information, on a somewhat experimental basis. It would be neat to be able to e.g. exclude cloud services or just e.g. cloudflare from the search results.	2023-12-17 15:03:00 +01:00
Viktor Lofgren	bde68ba48b	Merge branch 'master' into asn-info	2023-12-17 14:00:23 +01:00
Viktor Lofgren	bf44805e69	(*) Rename EdgeDomain$domain into topDomain This variable had a very confusing name, and was dangerously easy to use in the wrong place with the result of getting something that only works as expected half the time. Ideally this class needs an overhaul, the assumptions it makes about domain names aren't great.	2023-12-17 14:00:07 +01:00
Viktor Lofgren	edf9aa2c23	(*) Rename EdgeDomain$domain into topDomain This variable had a very confusing name, and was dangerously easy to use in the wrong place with the result of getting something that only works as expected half the time. Ideally this class needs an overhaul, the assumptions it makes about domain names aren't great.	2023-12-17 13:59:54 +01:00
Viktor Lofgren	4801c47273	(crawling-model) Fix bug where CrawledDocument.getDomain() trimmed www-prefixes This had the knock-on effect of breaking the anchor tag loading in the processor for a lot of domains, since they'd grab domains for the wrong domain name.	2023-12-17 13:53:31 +01:00
Viktor Lofgren	bcad6492d6	(sideloader) Fix integration problems with sideloaders In encyclopedia, add a class "mw-content-text" that the WikiSpecialization class is looking for during pruning to give the articles a more fair treatment. Also add generator keywords based on the generator type provided, to ensure that these documents show up in appropriate filters. Further, add a new document flag value 'Sideloaded' to be able to distinguish these entries.	2023-12-17 13:28:17 +01:00
Viktor Lofgren	5ab2a22e88	(search) Fix result count back down to 1 per domain	2023-12-17 13:14:23 +01:00
Viktor Lofgren	d7bd540683	(*) Replace the ip2location IP geolocation data with ASN information from apnic.net. Doesn't really make sense to use ip2location as a middle man for information that is already freely available...	2023-12-16 21:55:04 +01:00
dreimolo	62954f98de	adds xl to help output	2023-12-16 19:41:41 +01:00
Viktor Lofgren	722b56c8ca	(index) Fix rare bug in the index-switching logic This is caused by a resource contention with the query code. The proper way to fix this is to use some form of synchronization, but that will slow the code down. So we just hammer it a few times and let the GC deal with the problem if it fails. Not optimal, but fast.	2023-12-16 18:57:35 +01:00
Viktor Lofgren	f3f12058dc	(assistant) Fix logic error in filtering related domains	2023-12-16 18:45:53 +01:00
Viktor Lofgren	3da38d0483	(assistant) Fix logic error in filtering related domains	2023-12-16 18:44:25 +01:00
Viktor Lofgren	d715b1f9ca	(search) Improve error handling in search parameters parsing The code now intercepts and deals with potential exceptions during the parsing of search parameters. This is in response to constant bad requests from bots which were cluttering the logs. A catch clause is added that suppresses these errors and redirects to the base URL.	2023-12-16 18:42:13 +01:00
Viktor Lofgren	e13fa25e11	(assistant) Clean up the site info related domains view by filtering viable domains	2023-12-16 18:37:09 +01:00
Viktor Lofgren	34d4834ff6	(assistant) Clean up the site info related domains view by filtering viable domains	2023-12-16 18:27:24 +01:00
Viktor Lofgren	117ddd17d7	(assistant) Fix bugs in IP flag emoji generation	2023-12-16 17:07:17 +01:00
Viktor Lofgren	6f2bf38f0e	(index) Fix off-by-1 error in the domain count limiter	2023-12-16 16:57:05 +01:00
Viktor Lofgren	320882c34a	(site-info) Try to discover the schema of the website with a site:-query The site info view can't blindly assume that every website supports https. To figure out which schema to use when linking to a site, execute a single-result search for site:domain.name and then grab the schema off the result. To allow this, a count parameter is introduced to doSiteSearch() in SearchOperator.	2023-12-16 16:34:53 +01:00
Viktor	8bbb533c9a	Merge pull request #62 from MarginaliaSearch/warc (WIP) Use WARCs in the crawler	2023-12-16 16:02:46 +01:00
Viktor Lofgren	3113b5a551	(warc) Filter WarcResponses based on X-Robots-Tags There really is no fantastic place to put this logic, but we need to remove entries with an X-Robots-Tags header where that header indicates it doesn't want to be crawled by Marginalia.	2023-12-16 15:58:27 +01:00
dreimolo	c0cc05177f	corrects protobuf.plugins.grpc	2023-12-16 14:24:41 +01:00
dreimolo	0b34d43804	workaround for failing mac on apple silicon deps	2023-12-16 14:22:11 +01:00
dreimolo	6c7d7427bf	Adds check for wget and curl, and valid sample archives	2023-12-16 14:14:58 +01:00
Viktor Lofgren	54ed3b86ba	(minor) Remove dead code.	2023-12-15 21:49:35 +01:00
Viktor Lofgren	2001d0f707	(converter) Add @Deprecated annotation to a few fields that should no longer be used.	2023-12-15 21:42:00 +01:00
Viktor Lofgren	0f9cd9c87d	(warc) More accurate filering of advisory records Further create records for resources that were blocked due to robots.txt; as well as tests to verify this happens.	2023-12-15 21:37:02 +01:00
Viktor Lofgren	2e7db61808	(warc) More accurate filering of advisory records We want to mute some of these records so that they don't produce documents, but in some cases we want a document to be produced for accounting purposes. Added improved tests that reach for known resources on www.marginalia.nu to test the behavior when encountering bad content type and 404s. The commit also adds some safety try-catch:es around the charset handling, as it may sometimes explode when fed incorrect data, and we do be guessing...	2023-12-15 21:31:16 +01:00
Viktor Lofgren	5329968155	(crawler) Update CrawlingThenConvertingIntegrationTest This commit updates CrawlingThenConvertingIntegrationTest with additional tests for invalid, redirecting, and blocked domains. Improvements have also been made to filter out irrelevant entries in ParquetSerializableCrawlDataStream.	2023-12-15 21:04:06 +01:00
Viktor Lofgren	2e536e3141	(crawler) Add timestamp to CrawledDocument records This update includes the addition of timestamps to the parquet format for crawl data, as extracted from the Warc stream. The parquet format stores the timestamp as a 64 bit long, seconds since unix epoch, without a logical type. This is to avoid having to do format conversions when writing and reading the data. This parquet field populates the timestamp field in CrawledDocument.	2023-12-15 20:23:27 +01:00
Viktor Lofgren	cf935a5331	(converter) Read cookie information Add an optional new field to CrawledDocument containing information about whether the domain has cookies. This was previously on the CrawledDomain object, but since the WarcFormat requires us to write a WarcInfo object at the start of a crawl rather than at the end, this information is unobtainable when creating the CrawledDomain object. Also fix a bug in the deduplication logic in the DomainProcessor class that caused a test to break.	2023-12-15 18:09:53 +01:00
Viktor Lofgren	fa81e5b8ee	(warc) Use a non-standard WARC header to convey information about whether a website uses cookies This information is then propagated to the parquet file as a boolean. For documents that are copied from the reference, use whatever value we last saw. This isn't 100% deterministic and may result in false negatives, but permits websites that used cookies but have stopped to repent and have the change reflect in the search engine more quickly.	2023-12-15 16:37:53 +01:00

1 2 3 4 5 ...

1497 Commits