QDG Knowledge Base Read-only viewer QWebHub
workflow

Scraper & File Generator Process

Version 1 · Create initial deep-dive process doc for the PunterScraper -> S3 flag file -> TroyenDataFilesGenerator pipeline, for debugging stuck/missing meetings

Scraper & File Generator Process

How the core pipeline actually works, step by step: PunterScraper scrapes punters.com.au/racenet, writes Mongo + drops an S3 flag file; TroyenDataFilesGenerator polls for that flag file and builds the RaceCard/DataDump JSON, LLM comments, and silks. Written for debugging a stuck or missing meeting — see Overview for the wider system, TroyenDataHelpers Reference for the admin API that sits on top of this data.

PunterScraper

Entry point & run shape

Console app, PunterScraper/Program.cs (System.CommandLine). Key flags: --race-type/-r (T/G/H/A=all, default A), --start-date/-s, --end-date/-e (DD/MM/YYYY, default today → today+2), --country/-c (ISO2), --course/-co (comma-separated, starts-with match), -ic/--ignoreCahce (yes, the flag name has that typo — Y forces a re-fetch, see caching below). There are also ~10 debug-only flags (--health-check, --diagnose-*, --list-meetings, etc.) that never touch the production write path.

ExecuteScraping walks every date from endDate down to startDate; for each date it runs Horse.ExecuteHorseRace / GrayHound.ExecuteGrayHoundRace / Harness.ExecuteHarnessRaces for the requested discipline(s) — each is a full, independent run of that discipline for that one date. Horse.cs is the reference implementation; Harness.cs/GrayHound.cs mirror its shape with per-discipline differences noted below.

Per-discipline run, in order:

  1. Phase 1 — meeting list: GraphQL meetingsIndexByStartEndTime against https://api.punters.com.au/racing for the date, sport-filtered per discipline.
  2. Per meeting: resolve Course (name/country match, see lookup logic below), resolve cross-provider tabMeetingId, upsert the Meeting doc. Skip the whole meeting if existingMeeting.isLocked (an admin lock from TroyenDataHelpers.MeetingsController.save-meeting).
  3. Per race in the meeting (Parallel.ForEach, max 3 concurrent, up to 3 retries each): Phase 2 — race detail, GraphQL getEventById plus REST calls (getOverviewCacheable, getSectionals, getResultCard) and racenet.com.au form-history calls (fullFormsBySelectionIds, competitorBySlug) via a separate host/cookie system from the punters.com.au ones. Upserts Horses per runner, then the Race doc (embedding Runners).
  4. After all races: recompute average prize money, re-read the meeting's current fileGeneratedAt/isBypass/isTAB from Mongo before the final save (so this write can't clobber a concurrent admin edit or generator result), re-save the Meeting.
  5. Write the S3 flag file — this is the last thing that happens for a meeting.

Mongo writes and dedup

Everything goes through Common.MongoDbHelper.SaveOrGetBySlug<T> — the scraper never calls the Mongo driver directly. SaveOrGetBySlug itself: match by _id/Slug → if found and both old+new are isLocked, skip entirely (warn) → for each field listed in frozenFields, keep the old value → bump ModifiedDate → ReplaceOne + audit log (skipped for Logs/api_request_logs/datadumps/racecards) → if not found, assign new slug/id, InsertOne.

Collection Dedup key
meetings mDate + mDiscipline + mCourseId + internalReference (falls back to without internalReference); internalReference is a deterministic hash of the provider's own meeting id.
races Matched by race number within that meeting's existing races, not primarily by external event id (though the doc id is itself a deterministic hash of meetingId+eventNumber).
horses Deterministic id = hash of discipline+horseName+countryCode+foalYear (disambiguates same-named horses by foal year); falls back to hRacenetSlug+hName(+hFoalYear) for legacy records.
lookups Course-name/country → canonical Course.id reconciliation. Exactly one match → records the lookup once, idempotent. Zero or multiple matches → checks for an existing resolved lookup by (source, sourceReference, discipline); if still none, saves an UNMAPPED lookup and returns null — the meeting is skipped entirely, logged "Course not found in DB". It never guesses a placeholder course. If a meeting is silently missing, this is the first place to check — look for an UNMAPPED lookup for that course/country combo.

The S3 flag file

  • Written by Common/DataHelpers.CopyFlagFile(Meeting) — bucket is the hard-coded literal "troyon-watchdog" (not env-configurable on the writer side; the reader side does read it from an env var, defaulting to the same string, so they stay in sync in practice).
  • Key: Pending/flag_M_{yyyyMMddHHmmss}_{randomHex}.json.
  • Contents: Id (the Mongo Meeting.id), Name, UpdatedAt, FileGeneratedAt and IsBypassed (both copied from the meeting doc's current values at write time), Discipline, IsMeeting. It's a pointer + a cache-state snapshot, not the race data itself — the generator re-reads Mongo for the actual content.
  • Written once per meeting, after all its races are scraped — never per-race, never per-discipline-batch. A Race-level flag-file overload exists in code but the scrapers never call it; the only caller anywhere is an admin "regenerate" action in TroyenDataHelpers/MeetingsController.

"Skip if already scraped" caching

Confirmed S3-based, not local disk (Common/S3StorageHelper.ShouldDownloadFile, called before every raw-data fetch): lists S3 keys under {keyPrefix}/{filePrefix}-*, finds the newest matching file, and only re-downloads if none exists, it's under 2048 bytes, or it's older than maxAge (default 6 hours). -ic Y bypasses this and always re-fetches. The bucket for these raw scrape snapshots is punter-cache — distinct from troyen-data (final RaceCard/DataDump output) and troyon-watchdog (flag files). A separate, unrelated local-disk cache (Common/CacheHelper.cs, hour-based expiry) exists for small lookups (e.g. course-mapping XML) — don't confuse the two when debugging a "why didn't this re-scrape" question.

Anti-bot: the Broker

The component actually named "Broker" in code (PunterScraper/Broker/server.js, a Node/Express relay — not a message queue) replays outgoing requests through curl-impersonate to match a real browser's TLS/JA3 fingerprint, optionally through a proxy and/or a headless-Chromium JS-execution pass. On the C# side (Common/DownloadHelpers.cs), broker use is on by default — every outgoing request is converted to a curl string and sent to a broker instead of being made directly. BROKER_URLS (env, comma-separated) lists broker instances (.env.sample shows two); the client health-checks and round-robins across them, retries up to 10 times per call with a randomized proxy/TLS-profile pick each attempt.

Note: ProxyRelay (the separate top-level Docker project — squid/privoxy fronting an authenticated upstream proxy) is not referenced by name anywhere in the scraper code; the Broker just takes a generic proxy address as a parameter, which may or may not point at a ProxyRelay instance. Don't assume a direct code-level link between the two beyond "the Broker can be configured to route through some proxy."

TroyenDataFilesGenerator

Trigger: polling, not events

Common/FileWatchDog runs a genuine poll loop — every WATCHDOG_POLL_SECONDS (default 60), it lists Pending/flag_*.json in the troyon-watchdog bucket and processes up to WATCHDOG_CONCURRENCY (default 3) concurrently. There is no S3 event notification or queue — pure polling. An embedded Kestrel API (port WATCHDOG_API_PORT, default 5100) exposes /processing-keys, /pending-keys, POST /set-priority, etc. for inspecting/reordering the queue only — none of these endpoints trigger generation.

On success, the flag file moves Pending/ → Done/ unconditionally, as long as the per-flag processing function doesn't throw — including when it caught its own internal errors and returned normally. A permanently-unparseable flag file still ends up in Done/; there's no dead-letter distinction between "processed" and "gave up".

Per-race build sequence

For each flag file: parse it (3 retries) → skip if FileGeneratedAt is recent and not bypassed (cache gate, default 3-hour window via CacheDurationHours) → resolve the race list from Mongo (all races for the meeting, or the one race) → load meeting/course/LLMSystemPrompt docs → per race, ordered by race number:

  1. Dispatch to Generators.Horse|Harness|Grayhound.BuildBaseRaceCard — this single call does all the S3 raw-data reads and silk generation for the race (comment in code: "build once").
  2. Save the DataDump (S3 bucket troyen-data + Mongo datadumps) — once, language-agnostic.
  3. For each of EN-AU / EN-UK / EN-US (see LLM section — these are regional styles, not languages): apply language-specific comments, save the RaceCard to S3 + Mongo racecards.
  4. UpdateFileGenTime(race) — sets fileGeneratedAt, isBypass=false, a runner-name hash.
  5. GenerateRsRaceCardsForRace — the RS-client variant step (below).
  6. Any exception in this per-race block is caught, logged, and the loop continues to the next race — a bad race never aborts the meeting, but it also never retries; it'll only regenerate if a new flag file arrives.

After all races: UpdateFileGenTime(meeting).

Per-discipline differences (not cosmetic)

  • Silk generation exists for Horse and Harness, not Greyhound. Horse/Harness check S3StorageHelper.FileExists on the troyen-silk bucket and call SilkLibrary.SilkGenerator.GenerateAndSaveSilkImage if missing. Greyhound just copies whatever hSilkURL is already on the Horses doc — no generation, no existence check.
  • Inconsistent singleton pattern for the silk generator: Horse uses a lazy, retry-on-failure singleton (explicitly to avoid "poisoning" the class if construction fails once); Harness uses a plain eager static readonly field — if its S3-resource-loading constructor ever fails transiently, that failure isn't self-healing the way Horse's is. Worth knowing if Harness silks mysteriously stop generating after a transient S3 blip.
  • BuildBaseRaceCard signatures differ: Horse takes no course param (always uses the meeting's display name); Harness/Grayhound take an optional Course and use its display name when available.
  • Each discipline reads a different set of raw S3 JSON files (form guide / overview / sectional / event-data keys differ) — the schemas are provider/discipline-specific, not shared.

LLM comment/spotlight generation

  • LlmTornado (Common/ContentGenerationHelper.cs) wrapping an OpenAI-compatible client; the generation call sites hard-code GPT-4o-mini (not read from config for this specific call).
  • System prompt comes from the Mongo LLMSystemPrompt collection (type comment_and_spot_light/_g/_h, IsSelected == true — same collection TroyenDataHelpers' SystemPromptsController manages). User message is the full DataDump JSON for the race.
  • Idempotency: computes SHA256(DataDump JSON + model name) and checks for an existing llmContent doc with that hash + raceId+language+style before calling the LLM at all.
  • "3 language variants" are actually 3 English regional styles — EN-AU/EN-UK/EN-US — not 3 languages. A single LLM call returns all three at once (parsed as a JSON object keyed by style); the loop over the three styles is sequential, but in practice only the first iteration calls the LLM — the other two find the already-saved llmContent and just split it out.
  • Failure mode: the whole comment-generation call is wrapped in one try/catch that swallows any exception and returns the race card unchanged — a timed-out/failed LLM call silently produces a card with no comments/spotlight, no retry, no error field on the card. The only way to notice is an absent Comment/Spotlight field or a log line.
  • Caveat found in the shared helper: its signature accepts cacheTheResponse/dataSlug/dataId params that are never used in the method body — the real caching happens one layer up in the generator, not inside the helper. If you're chasing a caching bug, don't assume that flag does anything by itself.

Silk generation logic (Horse/Harness only)

Horses.hSilkURL is a deterministic path the scraper already set (/{discipline}/{foalYear}/{slug}.png). The generator does a plain S3 existence check on that key in troyen-silk; if present, skip (log "Silk exists... skipping"); if not, render via SilkGenerator.GenerateAndSaveSilkImage and write the PNG to that same key. This check is purely S3-existence-based — it does not consult Horses.isSilkFinal/hSilkGeneratedAt (those flags instead gate whether the scraper overwrites the silk description on a later re-scrape — an upstream, unrelated concern).

RS-client race-card variants

A distinct step run after the main card/dump is fully saved: for each active client-config (Type=RaceCard, Source=RS) whose region matches the meeting's country, calls RsRaceCardHelper.TryBuildRaceCard, which reads the legacy read-only RS MySQL database directly (raw ADO.NET, no ORM). This requires Meeting.rsMeetingId to already be populated — that field is filled by a completely separate cron job (TroyenService's SyncMeetingRsIdJob), not by this pipeline. If that job hasn't run yet for a meeting, RS card generation silently no-ops (logs "No RS race data found... skipping") — this is a common reason an RS-client card is missing while the main TD card exists fine. The whole step is wrapped so a failure here can never roll back or affect the already-completed main generation for that race — an explicit isolation-by-design comment in the code.

RsRaceCardHelper's MySQL connection string falls back to a hard-coded literal password in source if the env var isn't set — already flagged as a rotation candidate in the Overview.

How downstream consumers know data is ready

There's no new flag file and no event — completion is Mongo document state only. race.fileGeneratedAt/isBypass=false/hash and the equivalent meeting fields are the signal; TroyenAPI/TroyenDataHelpers read racecards/datadumps/meetings.fileGeneratedAt directly by polling the database, not by being pushed to. Moving the flag file to Done/ is bookkeeping for the watchdog's own queue only — nothing downstream reads it.

Error handling summary (no dead-letter, no alerting)

  • Flag-file parse failure → 3 retries, then logged and the file still moves to Done/ on the next cycle regardless.
  • Per-race failure (missing S3 source file, silk error, LLM error) → caught, logged (Console.WriteLine), loop continues to the next race. No retry, no requeue — that race only regenerates if a fresh flag file shows up later.
  • All error visibility is Console.WriteLine/NLog log lines — no structured error field on any document, no dead-letter queue, no alert hook.

Debugging checklist

  • Meeting never appears: check lookups for an UNMAPPED entry matching the course/country — the scraper skips meetings it can't resolve a course for, silently (logged, not erroring).
  • Meeting scraped but no RaceCard/DataDump: check whether the flag file actually landed in Pending//Done/ in troyon-watchdog, and check the generator's console/NLog output for a per-race exception around that time.
  • Card has no comments/spotlight: almost always a swallowed LLM exception — check logs for the printed exception around generation time; there's no other trace of it.
  • RS-client card missing but TD card present: check whether Meeting.rsMeetingId is populated yet — if TroyenService's RS-ID sync job hasn't run for that meeting, RS card generation no-ops by design.
  • Silks not generating for Harness specifically, after being fine for a while: consider the eager-static SilkGenerator construction in Harness.cs — a one-time transient S3 failure there can permanently disable silk generation for that discipline until the process restarts, unlike Horse's self-healing lazy singleton.
Updated by Claude on Aug. 11, 2026, 10:14 a.m. · Task: updatewiki create scraper and filegenerator process doc