FormCrawler Overview
Version 3 · Added Racing Australia weekly-index range discovery and launcher to the overview and updated verification to 75 tests.
Historical versionFormCrawler
Last reviewed: 2026-08-08 AEST
Purpose
FormCrawler is the policy-controlled acquisition component of the racing-data ingestion workflow. It discovers in-scope source URLs, retrieves public source pages or first-party documents, and preserves immutable artifacts with metadata for FormParser.
FormCrawler owns acquisition, request scheduling, retry state, rate limiting, diagnostics, immutable storage, and FormParser inbox handoffs. It does not map source fields into canonical racing JSON or update the racing database; those responsibilities belong to FormParser and downstream systems.
Implemented country modules
| Country | Source ID | Main capture forms | Country page |
|---|---|---|---|
| Australia | au_racing_australia |
Weekly meeting-index HTML plus five meeting lifecycle HTML pages | country-australia-racing-australia |
| Germany | de_deutscher_galopp |
Post-race race HTML | country-germany-deutscher-galopp |
| Chile | cl_hipodromo_chile |
Meeting/race/workout JSON | country-chile-hipodromo-chile |
| Uruguay | uy_maronas |
Calendar JSON, meeting XML, horse JSON | country-uruguay-maronas |
| Argentina | ar_stud_book |
Monthly calendar and complete meeting HTML | country-argentina-stud-book |
The form-crawler-programs page is the command and module index. Each country page documents its source routes, source data structure, launcher and generic commands, readiness rules, verification, limitations, and module-specific changelog.
Architecture
Calendar, meeting, race, horse, or supplementary source
|
v
Country adapter
|
v
Persistent target queue -> policy-controlled fetch -> source validation
|
v
immutable artifact
|
v
FormParser handoff
Reusable layers remain independent of country rules:
policy.pyandhttp.pyenforce request safeguards and optional proxy transport.storage.pywrites immutable artifacts and URI-based handoffs.state/base.pydefines backend-neutral crawl state.state/sqlite.pysupplies the local development queue and capture ledger.workflow.pycoordinates discovery, date ranges, worker leases, retries, and fan-out.sources/<country>/...contains strict source URL, discovery, and readiness rules.cli.pyexposes direct capture, discovery, queue, meeting, exact-discovery, date-range, and status commands.
Data contract and storage
Program source remains in C:\Users\Robert\Projects\FormCrawler. Runtime data remains outside the repository under C:\Users\Robert\project-data locally or the configured FORMCRAWLER_DATA_ROOT; Linux containers default to /data.
FormCrawler\t-artifacts\<year>\<race-code>\<source-id>\<event>\
FormParser\inbox\T\<year>\<race-code>\<source-id>\<artifact-id>.json
Implemented racing-code/ISO pairs are AUS/AU, GER/DE, CHI/CL, URU/UY, and ARG/AR. Metadata and handoffs identify source URL, retrieval time, content type/charset, phase, page and artifact type, source keys, content hash, and backend-neutral content/metadata URIs.
Direct responses remain byte-for-byte immutable. Explicitly proxied html_data is stored as UTF-8 with racingandsports_proxy provenance.
Friendly scraping policy
All source requests share central safeguards:
- one concurrent request per host;
- 10-second base delay plus 5-15 seconds of jitter, giving a normal 15-25-second interval;
- bounded retries and backoff;
- 25 MB response limit;
- circuit breaking;
- persistent shared host timing;
- first 403/429 pause and second consecutive block hard stop;
- external
FormCrawler\STOPkill switch.
Robert's Racing and Sports gateway is an explicit public-GET transport, not an automatic bypass. Both gateway and destination hosts are rate-limited and the selected country adapter still validates returned content.
Operation
Country launchers:
.\scripts\crawl_racing_australia_range.bat <start> <end> [limit]
.\scripts\crawl_racing_australia_meeting.bat "<meeting-url>"
.\scripts\crawl_deutscher_galopp_results.bat
.\scripts\crawl_hipodromo_chile_results.bat <start> <end> [limit]
.\scripts\crawl_hipodromo_chile_workouts.bat
.\scripts\crawl_maronas_results.bat <start> <end> [limit]
.\scripts\crawl_maronas_horse_profile.bat "<horse-name>"
.\scripts\crawl_stud_book_argentina_results.bat <start> <end> [limit]
Generic queue check:
python scripts\formcrawler.py queue-status
Run tests:
$env:PYTHONPATH="C:\Users\Robert\Projects\FormCrawler\src"
python -W error::ResourceWarning -m unittest discover -s tests -v
Deployment
FormCrawler is required to run in Docker on Amazon EKS. Core workflows and source adapters depend on interfaces rather than Windows paths or SQLite behavior. The filesystem store, SQLite state, and local rate-limit coordinator are development adapters. Production still requires shared object storage, distributed queue/state, and distributed rate-limit implementations.
Verified status
As of 2026-08-08:
- all five country adapters have passed live acquisition checks;
- source-specific artifact hashes and FormParser handoffs have been verified for representative captures;
- the current project suite passes 75 tests with
ResourceWarningtreated as an error; - the initial local commit remains
816f590; later working-tree changes have not yet been committed or pushed.
Current limitations
- FormParser does not yet formally validate and consume the artifact/handoff contract.
- Retention and repeated-snapshot content-deduplication policy remains undefined.
- Conditional HTTP requests using stored
ETagandLast-Modifiedvalues are not implemented. - Scheduled discovery has not been assigned an operating timetable.
- EKS production adapters remain to be implemented.
- Stud Book Argentina historical month selection requires interactive verification; FormCrawler rejects the site's silent current-month fallback.