QDG Knowledge Base Read-only viewer QWebHub
general

Scraping URL's List

Version 9 · Added the verified pushed branch heads, test gates, and common-baseline warning.

Historical version

Scraping URL's List

<!-- repository-baseline:start -->

Repository baseline

  • Thoroughbred: origin/main at 4d68b42, 25 adapters, 158 passing tests.
  • Harness/Greyhound: origin/codex/harness-greyhound-jurisdictions at 490b0c9, eight adapters including LeTROT and separate T/H/G registries, 131 passing tests.
  • Both branches are pushed and clean, but are not merged. Check the parent revision before starting a new branch; there is not yet one revision containing all disciplines.
  • See overview for the complete country/source Pre/Post/Others matrix. <!-- repository-baseline:end -->

Purpose

This page distinguishes sources that FormCrawler is configured to access from candidate racing websites supplied for future discovery and capture work. Active sources include their executable identifiers, allowed hosts, supported entry points, and implementation status. Candidate sources are reference inventory only and are not approved crawl scope.

The source-adapter registry remains the executable authority. Update this page whenever an active source adapter, allowed host, control URL, supported page pattern, or candidate source catalogue entry is added, changed, disabled, or removed.

Active sites

Status Source Source ID Country Discipline Racing code Canonical site
Active Racing Australia au_racing_australia Australia (AU) Thoroughbred (T) AUS https://www.racingaustralia.horse/
Active; licensed transport pending Equibase USA us_equibase USA (US) Thoroughbred (T) USA https://www.equibase.com/
Active Harness Racing Australia au_harness_racing Australia (AU) Harness (H) AUS https://natsite.harness.org.au/
Active Harness Racing New Zealand nz_hrnz New Zealand (NZ) Harness (H) NZ https://infohorse.hrnz.co.nz/
Active Standardbred Canada ca_standardbred_canada Canada (CA) Harness (H) CAN https://standardbredcanada.ca/
Active Deutscher Galopp de_deutscher_galopp Germany (DE) Thoroughbred (T) GER https://www.deutscher-galopp.de/
Active HVT de_hvt_harness Germany (DE) Harness (H) GER https://www.hvtonline.de/
Active LeTROT fr_letrot France (FR) Harness (H) FRA https://www.letrot.com/
Active Hipódromo Chile cl_hipodromo_chile Chile (CL) Thoroughbred (T) CHI https://hipodromo.cl/
Active Maroñas uy_maronas Uruguay (UY) Thoroughbred (T) URU https://hipica.maronas.com.uy/
Active Stud Book Argentina ar_stud_book Argentina (AR) Thoroughbred (T) ARG https://www.studbook.org.ar/
Active SNAI Ippica it_snai_ippica_harness Italy (IT) Harness (H) ITA https://ippica.snai.it/
Active Greyhound Racing Ireland ie_gri Ireland (IE) Greyhound (G) IRE https://www.grireland.ie/
Active Kincsem Park Greyhound hu_kincsem_park_greyhound Hungary (HU) Greyhound (G) HUN https://mla.kincsempark.hu/racing-days/greyhound/

France LeTROT

Only public HTTPS on letrot.com and www.letrot.com is accepted and canonicalized to www.letrot.com. Supported controls are /courses/YYYY-MM-DD; supported captures are /courses/programme/YYYY-MM-DD/<meeting_id> for efficient whole-meeting lifecycle HTML and /courses/YYYY-MM-DD/<meeting_id>/<race_number> for explicit race-level capture. Qualification, replay, publication and official-programme links are documented but are outside the default crawl.

Deutscher Galopp

Allowed hosts

  • www.deutscher-galopp.de
  • deutscher-galopp.de

Only HTTPS is accepted. Valid URLs are canonicalized to www.deutscher-galopp.de; allowing a host does not permit unrestricted crawling of every path on that host.

Supported URL patterns

Page type URL or pattern What FormCrawler uses it for
Calendar/control page https://www.deutscher-galopp.de/gr/renntage/ Captures the racing calendar page and discovers in-scope race links when they are present.
Meeting page https://www.deutscher-galopp.de/gr/renntage/{meeting_id}/?d={YYYYMMDD} Captures one dated meeting page and discovers its individual race links. {meeting_id} must be numeric.
Individual race page https://www.deutscher-galopp.de/gr/renntage/rennen.php?id={race_id}&d={YYYYMMDD}&s=R Captures a pre-race or post-race race page. For the current post-race workflow, s=R is used and the page must show that a result is available before it is accepted.

Verified example

https://www.deutscher-galopp.de/gr/renntage/rennen.php?id=1364507&d=20260802&s=R

This seed URL was used on 2026-08-07 to discover all nine races at the 2026-08-02 meeting. The queue completed with nine captured targets and no failures.

Implementation reference

The registered adapter is src\formcrawler\sources\de\deutscher_galopp.py. It validates the HTTPS scheme, allowed hosts, supported paths, numeric identifiers, date format, and required query values before any capture is attempted.

<!-- usa-equibase-active:start -->

USA Equibase

Allowed host and controls

Only HTTPS on www.equibase.com is accepted. Supported controls are the entries index, summary-results index with SAP=TN, and full-chart PDF index with SAP=TN. Supported leaves are official USA complete-card HTML, complete-summary HTML, full/race chart PDFs, and published race GPS HTML under their exact static paths.

The adapter date-filters meeting links before queue insertion and does not invent GPS URLs. The source is implemented and licensed-browser validated, but unattended direct/proxy transport is currently challenged; production use requires a licensed non-interactive session/API transport. <!-- usa-equibase-active:end -->

<!-- candidate-site-inventory:start -->

Candidate source inventory

Last updated: 2026-08-07

Purpose

This page records candidate racing source websites for FormCrawler discovery and capture work. Inclusion does not mean that a source adapter exists or that the URL has been technically validated.

Sources and conventions

  • Base inventory: C:\Users\Robert\Downloads\Site Links.xlsx, supplied on 2026-08-07.
  • Denmark (DEN) and Morocco links: supplied directly by Robert on 2026-08-07.
  • Disciplines use the FormCrawler codes T (Thoroughbred), H (Harness), and G (Greyhound).
  • Proposed source IDs use lowercase ASCII snake case.
  • Reuse one source ID when the same website or feed appears in multiple jurisdictions or disciplines and one adapter can own its capture behaviour.
  • Split sources when they require independent adapters, allowed-host rules, authentication, or capture behaviour.

Source identity and directories

Every implemented input source needs a stable machine-readable source_id. It is the source-name segment in the artifact path:

<discipline-root>\<year>\<race_code>\<source_id>\<event_slug>-<event_date>\

The implemented German source is recorded as:

  • Germany: Deutscher Galopp (de_deutscher_galopp)
    • Directory: t-artifacts\2026\GER\de_deutscher_galopp\

In this example, GER is the controlled racing jurisdiction code and de_deutscher_galopp is the source_id. The identifier names the website, feed, or source definition—not an individual downloaded file.

The IDs below are proposed identifiers for planning. They become authoritative only when the corresponding source definition or adapter is implemented.

Thoroughbred (T)

Workbook jurisdictions without a supplied Thoroughbred URL: Australia (NSW), Australia (QLD), Australia (WA), Austria, Ireland, Peru, Panama, and Mexico.

Harness (H)

Workbook countries without a supplied Harness URL: Austria and Malta.

Greyhound (G)

Maintenance

When a source is implemented, confirm or revise its proposed source_id, then record its allowed hosts, discovery URLs, capture phases, artifact types, and page patterns in versioned non-secret FormCrawler configuration. Do not place credentials or session material in this page. <!-- candidate-site-inventory:end -->

Request controls applying to every configured site

Every discovery, pre-race, post-race, and supplementary request must pass through FormCrawler's central policy layer. Current defaults serialize requests per host and apply a 10-second base delay plus 5-15 seconds of jitter, giving a 15-25-second interval between same-host request starts.

The common policy also applies bounded retries and backoff, response-size limits, circuit breaking, duplicate URL removal, shared local rate-limit state, and the external STOP kill switch. Source-specific settings may be stricter; relaxing a default requires Robert's explicit approval.

Candidate sources are not configured

The candidate inventory above is not a list of permitted crawl targets. Promoting a candidate requires a separate source adapter or explicit source definition, a stable source ID, restricted hosts and paths, capture phase and artifact type, tests, and an update to the active-site section before live requests are enabled.

Maintenance checklist

When adding or changing a source:

  1. Record the organization, source ID, country, discipline, and racing code.
  2. List exact allowed hosts and require HTTPS unless an explicit exception is approved.
  3. List control URLs and supported page patterns; do not describe an entire domain as crawlable when only selected paths are supported.
  4. Record whether each endpoint is discovery, pre-race, post-race, or supplementary capture.
  5. Add adapter validation and discovery tests before making live requests.
  6. Confirm that all requests use the common queue and scraping policy.
  7. Update this page and the FormCrawler changelog with the delivered source capability.

Thoroughbred source research status - 2026-08-08

Every Thoroughbred URL supplied in documents/SITE_LINKS.md is now classified as implemented, partial after bounded validation, research-only, document-only, authentication-required, transport-blocked, or outside the Thoroughbred discipline. Jurisdictions without a supplied URL are explicitly marked unstarted. Canonical review links and example horse profiles are maintained in thoroughbred-source-research-2026-08 and the source-specific country pages.

Updated by Codex on Aug. 8, 2026, 5:23 a.m. · Task: FormCrawler complete baseline reconciliation 2026-08-08 · Commit: main:4d68b42;harness-greyhound:490b0c9