FormCrawler Dashboard
Version 1 · Create the operations dashboard user guide with commands, status semantics, storage, validation evidence, and deployment boundaries
FormCrawler Dashboard
Last reviewed: 2026-08-10 AEST
Purpose
The FormCrawler Dashboard starts and monitors daily raw-acquisition work across Thoroughbred, Harness, and Greyhound sources. Each source and lifecycle phase runs independently, so one blocked, unpublished, or failed source does not stop other jurisdictions.
The dashboard monitors queue and artifact metadata only. FormCrawler continues to preserve raw source responses; canonical racing-data parsing remains the responsibility of FormParser.
Current release
The first dashboard release is commit 3319e66 on the pushed codex/crawl-dashboard branch. It adds:
run-day, a concurrent daily process supervisor;dashboard, a local HTTP service and browser interface;- durable run and job status in SQLite;
- one independent subprocess, queue database, and log per source/phase job; and
- packaged HTML, CSS, and JavaScript assets with no additional runtime dependency.
This branch must be integrated with the current unified main baseline before it becomes the authoritative production starting point.
Prerequisites
- Python and the FormCrawler checkout.
- A writable external data root; the local default used during validation is
C:\Users\Robert\project-data. - Source adapters already registered in the appropriate Thoroughbred, Harness, or Greyhound module.
- Network access appropriate to each public source. Existing source policy, challenge handling, retry limits, host spacing, and the external kill switch remain active.
Start the dashboard
python scripts\formcrawler.py dashboard `
--host 127.0.0.1 `
--port 8765 `
--data-root "C:\Users\Robert\project-data"
Open http://127.0.0.1:8765/.
The service binds to loopback by default and has no remote authentication. Do not expose it on a LAN or public interface in its current form.
Start a daily run
Use the dashboard controls to choose the pre-race date, post-race date, and maximum number of worker processes. The same operation can be started directly:
python scripts\formcrawler.py run-day `
--pre-date YYYY-MM-DD `
--post-date YYYY-MM-DD `
--max-workers 8 `
--data-root "C:\Users\Robert\project-data"
Use repeated --source SOURCE_ID arguments for a subset. With no source filter, the planner considers every registered source and records unsupported phases as not_available rather than launching commands known not to work.
--max-workers accepts 1 through 32. Eight is the initial practical default. Scheduling is round-robin by source so the first worker wave reaches distinct sources where possible. The existing shared host gate still serializes and safely spaces requests to the same website, even when different source processes are active.
What the dashboard shows
- retained run history with previous/next navigation;
- filters for discipline and country/source;
- per-source pre-race and post-race job state;
- requested-day meeting and saved-file counts;
- queue progress, retries, failures, and unavailable phases; and
- bounded tails of each worker log.
The local JSON interface uses POST /api/runs to create a run and GET /api/snapshot to read current dashboard state.
Status meanings
| Status | Meaning |
|---|---|
queued |
The job is planned and waiting for a worker. |
running |
Its independent subprocess is active. |
completed |
The requested work completed productively. |
partial |
Some useful work completed, but targets remain unavailable, retryable, or failed. |
no_data |
Discovery succeeded and no meeting matched the requested day. |
failed |
The job ended without productive output because of an error. |
not_available |
The source does not expose that lifecycle in a form the unattended daily planner can run. |
A mixed productive run aggregates to partial; it is marked failed only when every runnable job fails.
Runtime data
Operational state is external to Git:
C:\Users\Robert\project-data\FormCrawler\runtime\dashboard\
operations.sqlite3
runs\<run-id>\<source-id>-<phase>.sqlite3
runs\<run-id>\<source-id>-<phase>.log
The shared database stores run and job summaries. Each source/phase database remains isolated from other workers. Raw artifacts and FormParser handoffs continue to use their discipline-specific t-artifacts, h-artifacts, and g-artifacts locations. The dashboard does not rewrite those artifacts or handoffs.
Source-planning exceptions
- Argentina, Chile, and Uruguay currently provide dated post-race acquisition but no unattended pre-race plan.
- German Galopp pre-race is unavailable to this planner, and post-race requires a known meeting or race seed; both are recorded without launching a known-failing dated command.
- Australia Harness, Italy SNAI, New Zealand HRNZ, and Ireland GRI use current-list fallbacks where a historical pre-race archive is not exposed.
- New Zealand results use the published current results list because filenames do not reliably identify a year.
- Standardbred Canada uses its dated entries index.
- German HVT uses the approved Racing and Sports proxy from the first request.
Source publication timing can legitimately produce partial or no_data; those states should not be relabelled as crawler defects without checking the source and worker log.
First full validation
The first eight-worker run used pre-race date 2026-08-10 and post-race date 2026-08-09. It produced:
- 73 requested-phase meeting captures;
- 91 files totalling 116,881,425 bytes;
- nine completed jobs, six partial, three no-data, four failed, and four not-available jobs; and
- zero integrity or handoff defects across all 91 files when SHA-256, byte length, metadata identity and URI, discipline/source fields, and FormParser handoffs were independently checked.
The visible failures were Racing Australia's prior-week index mismatch, an initial German Galopp dated-planning attempt now prevented by manual_seed planning, and two HRNZ connection timeouts. Partial jobs chiefly reflected data not yet published for the requested date.
The implementation compiled and the complete dashboard-branch suite passed 139 tests. Browser verification covered rendering, filtering, run-control enabled/disabled state, the worker-log dialog, and console errors.
Troubleshooting
- Check the selected run and source/phase status rather than only the aggregate status.
- Open the worker log from the dashboard and inspect its bounded tail.
- Treat
not_availableas a planning limitation andno_dataas a valid discovery outcome. - For
partial, verify whether the source had published every requested target before retrying. - For
failed, rerun only the affected source where practical and retain the existing run for comparison. - Check the external
FormCrawler\STOPfile and source-specific transport requirements before restarting blocked work.
Moving the service
The local implementation deliberately uses the Python standard library, local subprocesses, filesystem logs, and SQLite. A shared deployment should place the HTTP service behind authentication and replace the local operations store/process launcher with a shared database, scheduler, and object-store-compatible status adapter. The source command plans, artifact contracts, and dashboard JSON response can remain stable across that move.
Related project documentation
form-crawler-programs: command, module, source, and storage inventory.overview: authoritative source matrix and validation baseline.changelog: delivered changes and known limitations.- Repository guide:
documents\OPERATIONS_DASHBOARD.md.