Scraper & File Generator Process
Version 1 · Create initial deep-dive process doc for the PunterScraper -> S3 flag file -> TroyenDataFilesGenerator pipeline, for debugging stuck/missing meetings
Scraper & File Generator Process
How the core pipeline actually works, step by step: PunterScraper scrapes punters.com.au/racenet, writes Mongo + drops an S3 flag file; TroyenDataFilesGenerator polls for that flag file and builds the RaceCard/DataDump JSON, LLM comments, and silks. Written for debugging a stuck or missing meeting — see Overview for the wider system, TroyenDataHelpers Reference for the admin API that sits on top of this data.
PunterScraper
Entry point & run shape
Console app, PunterScraper/Program.cs (System.CommandLine). Key flags: --race-type/-r (T/G/H/A=all, default A), --start-date/-s, --end-date/-e (DD/MM/YYYY, default today → today+2), --country/-c (ISO2), --course/-co (comma-separated, starts-with match), -ic/--ignoreCahce (yes, the flag name has that typo — Y forces a re-fetch, see caching below). There are also ~10 debug-only flags (--health-check, --diagnose-*, --list-meetings, etc.) that never touch the production write path.
ExecuteScraping walks every date from endDate down to startDate; for each date it runs Horse.ExecuteHorseRace / GrayHound.ExecuteGrayHoundRace / Harness.ExecuteHarnessRaces for the requested discipline(s) — each is a full, independent run of that discipline for that one date. Horse.cs is the reference implementation; Harness.cs/GrayHound.cs mirror its shape with per-discipline differences noted below.
Per-discipline run, in order:
- Phase 1 — meeting list: GraphQL
meetingsIndexByStartEndTimeagainsthttps://api.punters.com.au/racingfor the date, sport-filtered per discipline. - Per meeting: resolve
Course(name/country match, see lookup logic below), resolve cross-providertabMeetingId, upsert theMeetingdoc. Skip the whole meeting ifexistingMeeting.isLocked(an admin lock fromTroyenDataHelpers.MeetingsController.save-meeting). - Per race in the meeting (
Parallel.ForEach, max 3 concurrent, up to 3 retries each): Phase 2 — race detail, GraphQLgetEventByIdplus REST calls (getOverviewCacheable,getSectionals,getResultCard) and racenet.com.au form-history calls (fullFormsBySelectionIds,competitorBySlug) via a separate host/cookie system from the punters.com.au ones. UpsertsHorsesper runner, then theRacedoc (embeddingRunners). - After all races: recompute average prize money, re-read the meeting's current
fileGeneratedAt/isBypass/isTABfrom Mongo before the final save (so this write can't clobber a concurrent admin edit or generator result), re-save theMeeting. - Write the S3 flag file — this is the last thing that happens for a meeting.
Mongo writes and dedup
Everything goes through Common.MongoDbHelper.SaveOrGetBySlug<T> — the scraper never calls the Mongo driver directly. SaveOrGetBySlug itself: match by _id/Slug → if found and both old+new are isLocked, skip entirely (warn) → for each field listed in frozenFields, keep the old value → bump ModifiedDate → ReplaceOne + audit log (skipped for Logs/api_request_logs/datadumps/racecards) → if not found, assign new slug/id, InsertOne.
| Collection | Dedup key |
|---|---|
meetings |
mDate + mDiscipline + mCourseId + internalReference (falls back to without internalReference); internalReference is a deterministic hash of the provider's own meeting id. |
races |
Matched by race number within that meeting's existing races, not primarily by external event id (though the doc id is itself a deterministic hash of meetingId+eventNumber). |
horses |
Deterministic id = hash of discipline+horseName+countryCode+foalYear (disambiguates same-named horses by foal year); falls back to hRacenetSlug+hName(+hFoalYear) for legacy records. |
lookups |
Course-name/country → canonical Course.id reconciliation. Exactly one match → records the lookup once, idempotent. Zero or multiple matches → checks for an existing resolved lookup by (source, sourceReference, discipline); if still none, saves an UNMAPPED lookup and returns null — the meeting is skipped entirely, logged "Course not found in DB". It never guesses a placeholder course. If a meeting is silently missing, this is the first place to check — look for an UNMAPPED lookup for that course/country combo. |
The S3 flag file
- Written by
Common/DataHelpers.CopyFlagFile(Meeting)— bucket is the hard-coded literal"troyon-watchdog"(not env-configurable on the writer side; the reader side does read it from an env var, defaulting to the same string, so they stay in sync in practice). - Key:
Pending/flag_M_{yyyyMMddHHmmss}_{randomHex}.json. - Contents:
Id(the MongoMeeting.id),Name,UpdatedAt,FileGeneratedAtandIsBypassed(both copied from the meeting doc's current values at write time),Discipline,IsMeeting. It's a pointer + a cache-state snapshot, not the race data itself — the generator re-reads Mongo for the actual content. - Written once per meeting, after all its races are scraped — never per-race, never per-discipline-batch. A
Race-level flag-file overload exists in code but the scrapers never call it; the only caller anywhere is an admin "regenerate" action inTroyenDataHelpers/MeetingsController.
"Skip if already scraped" caching
Confirmed S3-based, not local disk (Common/S3StorageHelper.ShouldDownloadFile, called before every raw-data fetch): lists S3 keys under {keyPrefix}/{filePrefix}-*, finds the newest matching file, and only re-downloads if none exists, it's under 2048 bytes, or it's older than maxAge (default 6 hours). -ic Y bypasses this and always re-fetches. The bucket for these raw scrape snapshots is punter-cache — distinct from troyen-data (final RaceCard/DataDump output) and troyon-watchdog (flag files). A separate, unrelated local-disk cache (Common/CacheHelper.cs, hour-based expiry) exists for small lookups (e.g. course-mapping XML) — don't confuse the two when debugging a "why didn't this re-scrape" question.
Anti-bot: the Broker
The component actually named "Broker" in code (PunterScraper/Broker/server.js, a Node/Express relay — not a message queue) replays outgoing requests through curl-impersonate to match a real browser's TLS/JA3 fingerprint, optionally through a proxy and/or a headless-Chromium JS-execution pass. On the C# side (Common/DownloadHelpers.cs), broker use is on by default — every outgoing request is converted to a curl string and sent to a broker instead of being made directly. BROKER_URLS (env, comma-separated) lists broker instances (.env.sample shows two); the client health-checks and round-robins across them, retries up to 10 times per call with a randomized proxy/TLS-profile pick each attempt.
Note: ProxyRelay (the separate top-level Docker project — squid/privoxy fronting an authenticated upstream proxy) is not referenced by name anywhere in the scraper code; the Broker just takes a generic proxy address as a parameter, which may or may not point at a ProxyRelay instance. Don't assume a direct code-level link between the two beyond "the Broker can be configured to route through some proxy."
TroyenDataFilesGenerator
Trigger: polling, not events
Common/FileWatchDog runs a genuine poll loop — every WATCHDOG_POLL_SECONDS (default 60), it lists Pending/flag_*.json in the troyon-watchdog bucket and processes up to WATCHDOG_CONCURRENCY (default 3) concurrently. There is no S3 event notification or queue — pure polling. An embedded Kestrel API (port WATCHDOG_API_PORT, default 5100) exposes /processing-keys, /pending-keys, POST /set-priority, etc. for inspecting/reordering the queue only — none of these endpoints trigger generation.
On success, the flag file moves Pending/ → Done/ unconditionally, as long as the per-flag processing function doesn't throw — including when it caught its own internal errors and returned normally. A permanently-unparseable flag file still ends up in Done/; there's no dead-letter distinction between "processed" and "gave up".
Per-race build sequence
For each flag file: parse it (3 retries) → skip if FileGeneratedAt is recent and not bypassed (cache gate, default 3-hour window via CacheDurationHours) → resolve the race list from Mongo (all races for the meeting, or the one race) → load meeting/course/LLMSystemPrompt docs → per race, ordered by race number:
- Dispatch to
Generators.Horse|Harness|Grayhound.BuildBaseRaceCard— this single call does all the S3 raw-data reads and silk generation for the race (comment in code: "build once"). - Save the
DataDump(S3 buckettroyen-data+ Mongodatadumps) — once, language-agnostic. - For each of
EN-AU/EN-UK/EN-US(see LLM section — these are regional styles, not languages): apply language-specific comments, save theRaceCardto S3 + Mongoracecards. UpdateFileGenTime(race)— setsfileGeneratedAt,isBypass=false, a runner-name hash.GenerateRsRaceCardsForRace— the RS-client variant step (below).- Any exception in this per-race block is caught, logged, and the loop continues to the next race — a bad race never aborts the meeting, but it also never retries; it'll only regenerate if a new flag file arrives.
After all races: UpdateFileGenTime(meeting).
Per-discipline differences (not cosmetic)
- Silk generation exists for Horse and Harness, not Greyhound. Horse/Harness check
S3StorageHelper.FileExistson thetroyen-silkbucket and callSilkLibrary.SilkGenerator.GenerateAndSaveSilkImageif missing. Greyhound just copies whateverhSilkURLis already on theHorsesdoc — no generation, no existence check. - Inconsistent singleton pattern for the silk generator: Horse uses a lazy, retry-on-failure singleton (explicitly to avoid "poisoning" the class if construction fails once); Harness uses a plain eager
static readonlyfield — if its S3-resource-loading constructor ever fails transiently, that failure isn't self-healing the way Horse's is. Worth knowing if Harness silks mysteriously stop generating after a transient S3 blip. BuildBaseRaceCardsignatures differ: Horse takes no course param (always uses the meeting's display name); Harness/Grayhound take an optionalCourseand use its display name when available.- Each discipline reads a different set of raw S3 JSON files (form guide / overview / sectional / event-data keys differ) — the schemas are provider/discipline-specific, not shared.
LLM comment/spotlight generation
LlmTornado(Common/ContentGenerationHelper.cs) wrapping an OpenAI-compatible client; the generation call sites hard-codeGPT-4o-mini(not read from config for this specific call).- System prompt comes from the Mongo
LLMSystemPromptcollection (typecomment_and_spot_light/_g/_h,IsSelected == true— same collection TroyenDataHelpers'SystemPromptsControllermanages). User message is the fullDataDumpJSON for the race. - Idempotency: computes
SHA256(DataDump JSON + model name)and checks for an existingllmContentdoc with that hash +raceId+language+stylebefore calling the LLM at all. - "3 language variants" are actually 3 English regional styles —
EN-AU/EN-UK/EN-US— not 3 languages. A single LLM call returns all three at once (parsed as a JSON object keyed by style); the loop over the three styles is sequential, but in practice only the first iteration calls the LLM — the other two find the already-savedllmContentand just split it out. - Failure mode: the whole comment-generation call is wrapped in one try/catch that swallows any exception and returns the race card unchanged — a timed-out/failed LLM call silently produces a card with no comments/spotlight, no retry, no error field on the card. The only way to notice is an absent
Comment/Spotlightfield or a log line. - Caveat found in the shared helper: its signature accepts
cacheTheResponse/dataSlug/dataIdparams that are never used in the method body — the real caching happens one layer up in the generator, not inside the helper. If you're chasing a caching bug, don't assume that flag does anything by itself.
Silk generation logic (Horse/Harness only)
Horses.hSilkURL is a deterministic path the scraper already set (/{discipline}/{foalYear}/{slug}.png). The generator does a plain S3 existence check on that key in troyen-silk; if present, skip (log "Silk exists... skipping"); if not, render via SilkGenerator.GenerateAndSaveSilkImage and write the PNG to that same key. This check is purely S3-existence-based — it does not consult Horses.isSilkFinal/hSilkGeneratedAt (those flags instead gate whether the scraper overwrites the silk description on a later re-scrape — an upstream, unrelated concern).
RS-client race-card variants
A distinct step run after the main card/dump is fully saved: for each active client-config (Type=RaceCard, Source=RS) whose region matches the meeting's country, calls RsRaceCardHelper.TryBuildRaceCard, which reads the legacy read-only RS MySQL database directly (raw ADO.NET, no ORM). This requires Meeting.rsMeetingId to already be populated — that field is filled by a completely separate cron job (TroyenService's SyncMeetingRsIdJob), not by this pipeline. If that job hasn't run yet for a meeting, RS card generation silently no-ops (logs "No RS race data found... skipping") — this is a common reason an RS-client card is missing while the main TD card exists fine. The whole step is wrapped so a failure here can never roll back or affect the already-completed main generation for that race — an explicit isolation-by-design comment in the code.
RsRaceCardHelper's MySQL connection string falls back to a hard-coded literal password in source if the env var isn't set — already flagged as a rotation candidate in the Overview.
How downstream consumers know data is ready
There's no new flag file and no event — completion is Mongo document state only. race.fileGeneratedAt/isBypass=false/hash and the equivalent meeting fields are the signal; TroyenAPI/TroyenDataHelpers read racecards/datadumps/meetings.fileGeneratedAt directly by polling the database, not by being pushed to. Moving the flag file to Done/ is bookkeeping for the watchdog's own queue only — nothing downstream reads it.
Error handling summary (no dead-letter, no alerting)
- Flag-file parse failure → 3 retries, then logged and the file still moves to
Done/on the next cycle regardless. - Per-race failure (missing S3 source file, silk error, LLM error) → caught, logged (
Console.WriteLine), loop continues to the next race. No retry, no requeue — that race only regenerates if a fresh flag file shows up later. - All error visibility is
Console.WriteLine/NLog log lines — no structured error field on any document, no dead-letter queue, no alert hook.
Debugging checklist
- Meeting never appears: check
lookupsfor anUNMAPPEDentry matching the course/country — the scraper skips meetings it can't resolve a course for, silently (logged, not erroring). - Meeting scraped but no RaceCard/DataDump: check whether the flag file actually landed in
Pending//Done/introyon-watchdog, and check the generator's console/NLog output for a per-race exception around that time. - Card has no comments/spotlight: almost always a swallowed LLM exception — check logs for the printed exception around generation time; there's no other trace of it.
- RS-client card missing but TD card present: check whether
Meeting.rsMeetingIdis populated yet — ifTroyenService's RS-ID sync job hasn't run for that meeting, RS card generation no-ops by design. - Silks not generating for Harness specifically, after being fine for a while: consider the eager-static
SilkGeneratorconstruction inHarness.cs— a one-time transient S3 failure there can permanently disable silk generation for that discipline until the process restarts, unlike Horse's self-healing lazy singleton.