Scraping URL's List
Version 1 · Created the maintained scraping-site inventory with the active Deutscher Galopp hosts, supported URL patterns, request controls, and source-addition checklist.
Historical versionScraping URL's List
Purpose
This page is the maintained inventory of external sites and URL patterns that FormCrawler is configured to access. It records the source identifier, allowed hosts, supported entry points, page types, and current implementation status.
The source-adapter registry remains the executable authority. This page must be updated whenever a source adapter, allowed host, control URL, or supported page pattern is added, changed, disabled, or removed.
Active sites
| Status | Source | Source ID | Country | Discipline | Racing code | Canonical site |
|---|---|---|---|---|---|---|
| Active | Deutscher Galopp | de_deutscher_galopp |
Germany (DE) |
Thoroughbred (T) |
GER |
https://www.deutscher-galopp.de/ |
Deutscher Galopp
Allowed hosts
www.deutscher-galopp.dedeutscher-galopp.de
Only HTTPS is accepted. Valid URLs are canonicalized to www.deutscher-galopp.de; allowing a host does not permit unrestricted crawling of every path on that host.
Supported URL patterns
| Page type | URL or pattern | What FormCrawler uses it for |
|---|---|---|
| Calendar/control page | https://www.deutscher-galopp.de/gr/renntage/ |
Captures the racing calendar page and discovers in-scope race links when they are present. |
| Meeting page | https://www.deutscher-galopp.de/gr/renntage/{meeting_id}/?d={YYYYMMDD} |
Captures one dated meeting page and discovers its individual race links. {meeting_id} must be numeric. |
| Individual race page | https://www.deutscher-galopp.de/gr/renntage/rennen.php?id={race_id}&d={YYYYMMDD}&s=R |
Captures a pre-race or post-race race page. For the current post-race workflow, s=R is used and the page must show that a result is available before it is accepted. |
Verified example
https://www.deutscher-galopp.de/gr/renntage/rennen.php?id=1364507&d=20260802&s=R
This seed URL was used on 2026-08-07 to discover all nine races at the 2026-08-02 meeting. The queue completed with nine captured targets and no failures.
Implementation reference
The registered adapter is src\formcrawler\sources\de\deutscher_galopp.py. It validates the HTTPS scheme, allowed hosts, supported paths, numeric identifiers, date format, and required query values before any capture is attempted.
Request controls applying to every listed site
Every discovery, pre-race, post-race, and supplementary request must pass through FormCrawler's central policy layer. Current defaults serialize requests per host and apply a 10-second base delay plus 5-15 seconds of jitter, giving a 15-25-second interval between same-host request starts.
The common policy also applies bounded retries and backoff, response-size limits, circuit breaking, duplicate URL removal, shared local rate-limit state, and the external STOP kill switch. Source-specific settings may be stricter; relaxing a default requires Robert's explicit approval.
Sites not currently configured
No third-party sectional-time, analysis, ratings, or supplementary-data sites are currently registered. Adding one requires a separate source adapter or explicit source definition, a stable source ID, a restricted host/path list, capture phase and artifact type, tests, and an update to this inventory before live requests are enabled.
Maintenance checklist
When adding or changing a source:
- Record the organization, source ID, country, discipline, and racing code.
- List exact allowed hosts and require HTTPS unless an explicit exception is approved.
- List control URLs and supported page patterns; do not describe an entire domain as crawlable when only selected paths are supported.
- Record whether each endpoint is discovery, pre-race, post-race, or supplementary capture.
- Add adapter validation and discovery tests before making live requests.
- Confirm that all requests use the common queue and scraping policy.
- Update this page and the FormCrawler changelog with the delivered source capability.