FormCrawler Overview
Version 1 · Created the initial FormCrawler overview covering scope, architecture, storage, scraping safeguards, operation, deployment constraints, and verified status.
Historical versionFormCrawler
Purpose
FormCrawler is the acquisition component of the racing-data ingestion workflow. It discovers in-scope source URLs, retrieves source pages under a central scraping policy, and preserves immutable source artifacts with retrieval metadata for FormParser.
The first implemented source is Deutscher Galopp post-race data. Starting from one race URL, FormCrawler can discover the other races at the meeting, persist the targets, and capture them sequentially.
Responsibility boundary
FormCrawler owns:
- control-page and race-link discovery;
- request scheduling, retry state, rate limiting, and acquisition diagnostics;
- immutable raw response storage and content hashes;
- retrieval metadata and FormParser inbox handoffs.
FormCrawler does not parse racing values, map fields into a canonical schema, calculate ratings, or update the core racing database. Those responsibilities belong to FormParser and downstream systems. Supplementary source material such as sectional-time analysis may be captured, but derived or third-party calculated data must remain distinct from core source artifacts.
Architecture
Calendar, meeting, or race URL
|
v
Source adapter
|
v
Persistent target queue -> policy-controlled fetch -> immutable artifact
|
v
FormParser handoff
Reusable modules are separated from country-specific rules:
policy.pyandhttp.pyenforce shared acquisition safeguards.storage.pywrites immutable artifacts and URI-based handoffs.state/base.pydefines the backend-neutral crawl-state contract.state/sqlite.pysupplies the local development queue and capture ledger.workflow.pycoordinates discovery, worker leases, retries, and meeting fan-out.sources/de/deutscher_galopp.pycontains only Deutscher Galopp URL, discovery, and capture-readiness rules.cli.pyexposes direct capture, discovery, queue processing, meeting capture, and queue status commands.
Data contract and storage
Program source remains in C:\Users\Robert\Projects\FormCrawler. Runtime data must remain outside the repository under C:\Users\Robert\project-data locally or the configured FORMCRAWLER_DATA_ROOT.
Thoroughbred artifacts use this local structure:
FormCrawler\t-artifacts\<year>\<race_code>\<source_id>\<event_date>\<artifact_id>\
FormParser\inbox\T\<year>\<race_code>\<source_id>\<artifact_id>.json
Germany uses racing jurisdiction code GER, ISO country code DE, and discipline T. Metadata and handoffs use backend-neutral content and metadata URIs; the filesystem adapter emits file: URIs.
The local queue database defaults to:
C:\Users\Robert\project-data\FormCrawler\runtime\queue\formcrawler.sqlite3
The queue stores targets, discovery provenance, capture attempts, artifact history, retry state, and expiring worker leases. Targets are deduplicated by source, canonical URL, capture phase, and artifact type.
Friendly scraping policy
All source requests pass through the common policy layer. Same-host requests are serialized and use a 10-second base delay plus 5-15 seconds of jitter, producing a 15-25-second interval. The implementation also includes bounded retries, backoff, response-size limits, circuit breaking, shared local rate-limit state, and a STOP kill switch under the external data root.
Operation
Capture a complete Deutscher Galopp meeting from one post-race URL:
python scripts\formcrawler.py capture-meeting `
--source de_deutscher_galopp `
--seed-race-url "https://www.deutscher-galopp.de/gr/renntage/rennen.php?id=1364507&d=20260802&s=R" `
--data-root "C:\Users\Robert\project-data"
Discovery and queue processing can also be separated:
python scripts\formcrawler.py discover --source de_deutscher_galopp --url <control-url>
python scripts\formcrawler.py run-queue --source de_deutscher_galopp
python scripts\formcrawler.py queue-status
Run the test suite:
$env:PYTHONPATH="C:\Users\Robert\Projects\FormCrawler\src"
python -m unittest discover -s tests -v
Deployment
FormCrawler is required to run in Docker on Amazon EKS. Core workflows and source adapters therefore depend on interfaces rather than Windows paths or SQLite behavior. The filesystem artifact store, SQLite state store, and local rate-limit coordination are development adapters. Production deployment still requires shared object storage, distributed queue/state, and distributed rate-limit implementations.
Container data configuration uses FORMCRAWLER_DATA_ROOT; Linux defaults to /data.
Verified status
On 2026-08-07:
- all Python modules compiled;
- 14 unit tests passed;
- the supplied Deutscher Galopp seed race discovered all 9 races at the meeting;
- the queue completed with 9 captured and 0 failed;
- artifact SHA-256 values and FormParser URI references were verified;
- eight observed same-host request gaps ranged from 15.949 to 24.660 seconds.
The initial implementation is committed locally as 816f590 (Add German post-race meeting acquisition workflow).
Current limitations
- FormParser does not yet formally validate and consume the
1.0.0handoff contract. - Retention and content-hash deduplication policy for repeated snapshots remains to be defined.
- Conditional HTTP requests using stored
ETagandLast-Modifiedvalues are not implemented. - Scheduled calendar discovery has not been assigned an operating schedule.
- EKS production adapters are not yet implemented.