CatgirlIntelligenceAgency

Author	SHA1	Message	Date
Viktor Lofgren	b74a3ebd85	(crawler) WIP integration of WARC files into the crawler process. At this stage, the crawler will use the WARCs to resume a crawl if it terminates incorrectly. This is a WIP commit, since the warc files are not fully incorporated into the work flow, they are deleted after the domain is crawled. The commit also includes fairly invasive refactoring of the crawler classes, to accomplish better separation of concerns.	2023-12-11 19:32:58 +01:00
Viktor Lofgren	45987a1d98	Merge branch 'master' into warc	2023-12-11 14:32:35 +01:00
Viktor Lofgren	8ef34883a8	(search) Move site information out of the search service and into assistant. This reduces the impact of restarting the search service, as the site information takes a few minutes to load during which it's not available. It also permits exposing this information via API in the future if there is interest in this. The assistant service was also modified to do a late load of the suggestions trie, as this is a major contributor to its start-up time. Finally, some changes were made to the client library, a new get() method was added that takes a TypeToken to allow deserialization of generics such as List<Foo>, and the scheduler was also modified to use virtual threads.	2023-12-09 16:30:06 +01:00
Viktor Lofgren	072b5fcd12	Implement Warc-recording wrapper for OkHttp3 client This is a first step of using WARC as an intermediate flight recorder style step in the crawler, ultimately aimed at being able to resume crawls if the crawler is restarted. This component is currently not hooked into anything. The OkHttp3 client wrapper class 'WarcRecordingFetcherClient' was implemented for web archiving. This allows for the recording of HTTP requests and responses. New classes were introduced, 'WarcDigestBuilder', 'IpInterceptingNetworkInterceptor', and 'WarcProtocolReconstructor'. The JWarc dependency was added to the build.gradle file, and relevant unit tests were also introduced. Some HttpFetcher-adjacent structural changes were also done for better organization.	2023-12-08 13:49:16 +01:00
Viktor Lofgren	280132dad0	(search) Fix script loading for mobile support	2023-12-02 17:06:40 +01:00
Viktor Lofgren	7c8a60b8cf	(search) Site info view is mostly done Also optimize the rendering a bit to avoid having to allocate huge string buffers, writing directly to Spark's response instead.	2023-12-02 17:06:40 +01:00
Viktor Lofgren	a258f0af7a	(search) Refactor search parameters to include query	2023-12-02 17:06:40 +01:00
Viktor Lofgren	01621c6344	(renderer) Make helpers configurable on a by-service basis.	2023-12-02 17:06:40 +01:00
Viktor Lofgren	1dafa0c74d	(mqapi/control) Repair repartition endpoint, deprecate notify endpoints. The repartition endpoint was mis-addressing its mqapi notifications, omitting the proper nodeId. In fixing this, it became apparent that having both @MqRequest and @MqNotification is a serious footgun, and the two should be unified into a single API where the caller isn't burdened with knowledge of the remote end's implementation specifics.	2023-11-27 16:01:12 +01:00
Viktor Lofgren	dd507a3808	(db) Fix migrations, bump flyway to 10.0.1 Tricky problem, creating a procedure apparently needs delimiter shenanigans in Flyway, otherwise it will truncate the END statement and mariadb will be sad.	2023-11-21 20:04:35 +01:00
Viktor Lofgren	f58a9f46be	(loader) Don't truncate the entire links table on load This behavior is an old vestige from the days of only having a single loader process. We'd truncate the links table because doing inserts/updates was too slow. This was also important because we had 32 bit ID, and there's a lot of links between domains to go around... Instead we delete the rows associated with the current node with a stored procedure PURGE_LINKS_TABLE. We also update the PRIMARY KEY to a BIGINT. We'll need to load the data in excess of billion times to hit an ID rollover, so it'll be fine.	2023-11-16 10:30:12 +01:00
Viktor Lofgren	858357a246	(metrics) Get prometheus up out of disrepair * Fix bad labels * Add nodeId where appropriate * Hopefully fix histogram buckets for index query times	2023-11-08 14:01:28 +01:00
Viktor Lofgren	7aa2f80117	(domain) id.au should be treated as a TLD	2023-11-06 19:07:47 +01:00
Viktor Lofgren	2b77184281	(converter) Integrate atags with the topology field	2023-11-06 13:46:44 +01:00
Viktor Lofgren	0152004c42	Initial Commit Anchor Tags * Added new (optional) model file in $WMSA_HOME/data/atags.parquet * Converter gets a component for creating a projection of its domains onto the full atags parquet file * New WordFlag ExternalLink * These terms are also for now flagged as title words * Fixed a bug where Title words aliased with UrlDomain words * Fixed a bug in the encyclopedia sideloader that gave everything too high topology ranking	2023-11-04 14:24:17 +01:00
Viktor Lofgren	659743b39c	(executor) Export Data actor allocates its own storage	2023-10-31 17:04:07 +01:00
Viktor Lofgren	5d6e0e3790	(log) Clean up logging Don't log the PROCESS stream to executor's logs, as it will also be logged in the spawned process' log files. Also tell the spawned process which "service" it is so that it gets a log file with a name that makes sense.	2023-10-29 15:52:17 +01:00
Viktor Lofgren	0f637fb722	(logging) Better logging configurations	2023-10-26 12:48:10 +02:00
Viktor Lofgren	97fcbdd6d9	(control) Move storage actions into the actions tab * Also disable annoying CSS animations	2023-10-25 21:23:56 +02:00
Viktor Lofgren	d7686b665e	Refactoring * Encyclopedia sideloader; permit providing base URL. * Storage base shows node id in GUI * ProcessLivenessMonitorActor restarts automatically * Clean-up of outbox code	2023-10-25 18:51:02 +02:00
Viktor Lofgren	436a55ee1e	(control) Render UUID tooltip with dashes.	2023-10-24 16:37:40 +02:00
Viktor Lofgren	e4bddb4993	(control) Better UUID accessibility	2023-10-23 12:53:53 +02:00
Viktor Lofgren	758f9b5aa5	(converter) Get UUID pips out of the models Rendering concerns shouldn't be in the models, it's poor separation of concerns and very difficult to follow.	2023-10-22 14:24:52 +02:00
Viktor Lofgren	29ce8ca0cf	(db) Reduce db pool size This is a temporary thing	2023-10-22 14:03:09 +02:00
Viktor Lofgren	12fda1a36b	(control) Temporarily re-writing the data balancer to get it to work in prod Need to clean this up later.	2023-10-22 14:03:09 +02:00
Viktor Lofgren	c6abcd91fa	(control) Better use of FS states, fix bug with start/stop actors	2023-10-20 16:37:49 +02:00
Viktor Lofgren	d76d926c38	(control/executor) Add new configuration options for node It's now possible to configure prod instance to not retain processed data.	2023-10-20 14:05:19 +02:00
Viktor Lofgren	2b3c167845	(controller) Additional configuration options for node	2023-10-20 13:13:36 +02:00
Viktor Lofgren	584bb3a648	(fs) interface cleanup	2023-10-20 12:24:18 +02:00
Viktor Lofgren	23526f6d1a	(executor) Executor service now pulls DomainType list for CRAWL on "recrawl" This is an automatic integration with the submit-site repo on github and also crawl-queue.	2023-10-19 17:48:34 +02:00
Viktor Lofgren	23f0c79fba	(control) GUI for data sets/domain types.	2023-10-19 17:48:34 +02:00
Viktor Lofgren	81dd3809e9	(*) WIP Add node affinity to EC_DOMAIN Very messy commit due to fractalline yak shaving	2023-10-19 17:48:34 +02:00
Viktor Lofgren	84fea0fd05	(node) Nodes auto-start their monitor actors.	2023-10-16 15:33:22 +02:00
Viktor Lofgren	2df3e0f881	(node) Nodes auto-configure on start-up instead of requiring manual configuration.	2023-10-16 14:46:35 +02:00
Viktor Lofgren	39911e3acd	(control) Fix incorrect storage base and clean up GUI for data	2023-10-16 13:30:26 +02:00
Viktor Lofgren	3d1c15ef99	(client) Refactor liveness monitor	2023-10-16 12:34:01 +02:00
Viktor Lofgren	f718482e98	(client) Fix tests	2023-10-16 12:12:16 +02:00
Viktor Lofgren	8dafd13cd7	(client) Fix executor tests	2023-10-16 12:02:57 +02:00
Viktor Lofgren	0b19b28a64	(file-storage) Delete unused code	2023-10-16 12:02:57 +02:00
Viktor Lofgren	16e0738731	(*) Get multi-node routing working.	2023-10-15 18:38:30 +02:00
Viktor Lofgren	eacbf87979	(control) New list and form for index nodes.	2023-10-14 21:46:52 +02:00
Viktor Lofgren	108b4cb648	(service) Keep disabled multi-noded services dormant when they are configured to be disabled.	2023-10-14 20:58:55 +02:00
Viktor Lofgren	a9dff407a1	(config/db) Clean up migrations	2023-10-14 20:34:03 +02:00
Viktor Lofgren	6308a8dfcd	(control) Node configuration	2023-10-14 16:47:52 +02:00
Viktor Lofgren	4baf9527d7	() WIP Control GUI redesign, executor-service, multi-node mq This turned out to be very difficult to do in small isolated steps. Design overhaul of the control gui using bootstrap * Move the actors out of control-service into to a new executor-service, that can be run on multiple nodes * Add node-affinity to message queue	2023-10-14 12:08:43 +02:00
Viktor Lofgren	199c459697	(*) Add node-affinity to services, processes and file storage.	2023-10-10 12:32:22 +02:00
Viktor Lofgren	61288c5e68	(service, client) First steps towards multiple nodedness	2023-10-09 22:13:27 +02:00
Viktor Lofgren	3889c4bdd9	(refactor) Remove features-search and update documentation	2023-10-09 15:12:30 +02:00
Viktor Lofgren	89c6d85f2f	(query-service) Create new empty 'query-service' service	2023-10-08 17:31:50 +02:00
Viktor Lofgren	77ccab7d80	(index) Move linkdb to index from search. This makes index complete in the sense that you can deploy an index instance and build a complete separate application on top of it, without having to go through the Marginalia-laden search service.	2023-10-08 16:48:35 +02:00

1 2 3 4

196 Commits