QDG Knowledge Base Read-only viewer QWebHub
general

Scraping URL's List

Version 1 · Created the maintained scraping-site inventory with the active Deutscher Galopp hosts, supported URL patterns, request controls, and source-addition checklist.

Historical version

Scraping URL's List

Purpose

This page is the maintained inventory of external sites and URL patterns that FormCrawler is configured to access. It records the source identifier, allowed hosts, supported entry points, page types, and current implementation status.

The source-adapter registry remains the executable authority. This page must be updated whenever a source adapter, allowed host, control URL, or supported page pattern is added, changed, disabled, or removed.

Active sites

Status Source Source ID Country Discipline Racing code Canonical site
Active Deutscher Galopp de_deutscher_galopp Germany (DE) Thoroughbred (T) GER https://www.deutscher-galopp.de/

Deutscher Galopp

Allowed hosts

  • www.deutscher-galopp.de
  • deutscher-galopp.de

Only HTTPS is accepted. Valid URLs are canonicalized to www.deutscher-galopp.de; allowing a host does not permit unrestricted crawling of every path on that host.

Supported URL patterns

Page type URL or pattern What FormCrawler uses it for
Calendar/control page https://www.deutscher-galopp.de/gr/renntage/ Captures the racing calendar page and discovers in-scope race links when they are present.
Meeting page https://www.deutscher-galopp.de/gr/renntage/{meeting_id}/?d={YYYYMMDD} Captures one dated meeting page and discovers its individual race links. {meeting_id} must be numeric.
Individual race page https://www.deutscher-galopp.de/gr/renntage/rennen.php?id={race_id}&d={YYYYMMDD}&s=R Captures a pre-race or post-race race page. For the current post-race workflow, s=R is used and the page must show that a result is available before it is accepted.

Verified example

https://www.deutscher-galopp.de/gr/renntage/rennen.php?id=1364507&d=20260802&s=R

This seed URL was used on 2026-08-07 to discover all nine races at the 2026-08-02 meeting. The queue completed with nine captured targets and no failures.

Implementation reference

The registered adapter is src\formcrawler\sources\de\deutscher_galopp.py. It validates the HTTPS scheme, allowed hosts, supported paths, numeric identifiers, date format, and required query values before any capture is attempted.

Request controls applying to every listed site

Every discovery, pre-race, post-race, and supplementary request must pass through FormCrawler's central policy layer. Current defaults serialize requests per host and apply a 10-second base delay plus 5-15 seconds of jitter, giving a 15-25-second interval between same-host request starts.

The common policy also applies bounded retries and backoff, response-size limits, circuit breaking, duplicate URL removal, shared local rate-limit state, and the external STOP kill switch. Source-specific settings may be stricter; relaxing a default requires Robert's explicit approval.

Sites not currently configured

No third-party sectional-time, analysis, ratings, or supplementary-data sites are currently registered. Adding one requires a separate source adapter or explicit source definition, a stable source ID, a restricted host/path list, capture phase and artifact type, tests, and an update to this inventory before live requests are enabled.

Maintenance checklist

When adding or changing a source:

  1. Record the organization, source ID, country, discipline, and racing code.
  2. List exact allowed hosts and require HTTPS unless an explicit exception is approved.
  3. List control URLs and supported page patterns; do not describe an entire domain as crawlable when only selected paths are supported.
  4. Record whether each endpoint is discovery, pre-race, post-race, or supplementary capture.
  5. Add adapter validation and discovery tests before making live requests.
  6. Confirm that all requests use the common queue and scraping policy.
  7. Update this page and the FormCrawler changelog with the delivered source capability.
Updated by Codex on Aug. 7, 2026, 5:14 a.m. · Task: Create Scraping URL's List page · Commit: 816f590