QDG Knowledge Base Read-only viewer QWebHub
overview

Overview

Version 4 · Link the new Database Reference doc from Starting points

Historical version

Troyon Data — Racing Data Platform

Troyon Data ("Troyen"/"Troyon" are used interchangeably across the codebase — e.g. TroyenService builds as TroyonService) is a thoroughbred / harness / greyhound racing data pipeline and content platform.

  • Repository: https://github.com/TPL-Tech-Titans/Troyon-Data.git
  • Solution: ProjectTroyon.sln (only 12 of ~25 top-level folders are wired into it — the rest are standalone tools/scripts/legacy stubs)
  • .NET SDK: pinned to 8.0.0 via global.json (rollForward: latestMajor)

What it does

The system scrapes race-meeting data from bookmaker/aggregator sites — punters.com.au + racenet (PunterScraper), neds.com.au (NedsScraper) — and cross-references it against a legacy read-only MySQL system called "RS" (Racing & Sports). Everything is persisted in MongoDB (Troyon_Dev DB) as meetings / races (embedding runners) / horses documents — see the Database Reference for the full schema.

A second stage (TroyenDataFilesGenerator) watches S3 for "flag files" dropped by PunterScraper and turns raw meeting/race data into polished, LLM-enriched Race Cards and Data Dumps (JSON) — including AI-generated runner comments/spotlights (via LlmTornado/OpenAI, 3 language variants) and generated jockey silk colour images (SVG→PNG via SilkLibrary). Client-specific ("RS-client") race-card variants are also produced by reading the legacy RS MySQL directly. See the Scraper & File Generator Process doc for exactly how this works end to end.

Results are served externally via TroyenAPI (read-only; auth is a custom API-key scheme, not JWT — see TroyenAPI Reference) and internally via TroyenDataHelpers (read/write admin API, no enforced auth in code — see TroyenDataHelpers Reference), with TroyenDataWeb acting as the admin/ops portal for managing meetings/races/silks/comments. TroyenService runs three independent scheduled jobs (Windows Service, Cronos-scheduled) that enrich the same Mongo documents: RS-ID cross-reference, runner stats, scratchings.

Data flow

PunterScraper ──(S3 flag file: troyon-watchdog)──► TroyenDataFilesGenerator
  (punters.com.au +                                  (builds RaceCard/DataDump JSON,
   racenet, → MongoDB)                                 LLM comments, silks, RS-client cards)
                                                              │
NedsScraper ──(independent, own Web API)──► MongoDB ◄────────┘
  (neds.com.au)                                │
                                                ▼
                              TroyenAPI (external, read-only) ◄─┐
                              TroyenDataHelpers (internal, r/w) │
                                                │                │
                                                ▼                │
                                     TroyenDataWeb (admin portal, calls both)

TroyenService (independent cron jobs: RS-ID sync, runner stats, scratchings) ──► MongoDB

PunterScraper and TroyenDataFilesGenerator communicate only via the S3 flag file — there is no orchestrator or message queue, and MongoDB is the only write-side datastore in the core pipeline (RS MySQL is legacy and read-only).

Core projects (in ProjectTroyon.sln)

Project Responsibility Tech
Common Shared foundation: MongoDB access, S3 storage, LLM content generation, RS-MySQL race-card helper, scraper HTTP/broker helpers. Referenced by nearly everything. net8.0 lib — AWSSDK.S3, MongoDB.Driver, MySqlConnector, Selenium, LlmTornado
TroyenModels Shared BSON domain models (Meeting, Race, Horses, RaceCard, DataDump) used across the Mongo-backed pipeline. net8.0 lib — EF Core, MongoDB.Driver
TroyenData.Core Racing.PaceCalculator — effectively dead code, only referenced from commented-out scraper code. net8.0 lib, no packages
PunterScraper Production scraper: one-shot console/Docker job hitting punters.com.au GraphQL for meetings/races/horses/lookups, writes to Mongo, drops an S3 flag file. net8.0 console — refs Common, TroyenModels, TroyenData.Core. Docker: punter-scraper
NedsScraper Actively-maintained Web API scraper for neds.com.au; own changelog shows weekly development. Independent of the PunterScraper flow, writes the same Mongo collections. net8.0 Web API — HtmlAgilityPack, LlmTornado, Playwright, Serilog
SilkLibrary Renders jockey/horse "silk" colour SVGs to PNG (SkiaSharp), uploads to S3. net8.0 lib — SkiaSharp, Svg.Skia
TroyenDataFilesGenerator The "race card factory": polls S3 for PunterScraper flag files, builds RaceCard/DataDump JSON per discipline, runs LLM comment/spotlight generation, generates silks, builds RS-client variants. net8.0 Web/Exe. Docker: troyen-data-gen
TroyenAPI External, read-only customer-facing REST API for meetings/races/racecards. Custom API-key auth (not JWT), API versioning. IIS site TroyonAPI/TD-External. net8.0 Web, win-x64 self-contained. Docker: troyen-external-api
TroyenDataHelpers Internal read/write admin API (Meetings/RaceUpdate/Races/Horses controllers) — backend behind TroyenDataWeb. Bundles silk pattern SVGs. No enforced auth in code. IIS TroyenDataHelpersAPI/TD-Internal. net8.0 Web, win-x64. Docker: troyen-internal-api
TroyenDataWeb Internal Admin Portal: ASP.NET MVC, admin views for runners/races/comments/silks. Login is custom email + TOTP against MongoDB (AuthService/TotpService, not ASP.NET Identity/EF Core, despite EF Core/Pomelo.MySql packages being referenced in the project). Hosts install-guide docs for the standalone scraper tools. IIS TroyenDataWeb/TD-Web. net8.0 Web, win-x64. Docker: troyen-web
TroyenService Windows Service running 3 independent cron jobs against Mongo: RS-ID sync, runner stats, scratchings. Not an orchestrator — no reference to PunterScraper or TroyenDataFilesGenerator. net8.0 Worker — Cronos. Docker: troyen-services
TroyenRaceIngestor New, parallel ingestion pipeline (see below). net8.0 Web (embedded Kestrel queue API). Docker: troyen-race-ingestor, port 5101

Standalone tools / legacy (not in .sln)

  • PunterScraperLauncher — WPF Windows GUI wrapping a self-contained PunterScraper.exe build for non-technical staff; must be manually rebundled when PunterScraper/Common changes.
  • PunterWebScraper — DOM-based (HtmlAgilityPack) punters.com.au form-guide scraper, distributed as a VPN-gated zip; a plausible successor to a Selenium prototype but not wired into the main pipeline yet.
  • RacenetScraper — stale (~1 year); current Program.cs is a small Mongo batch query, not an active racenet.com.au scraper.
  • RacingPostscraper — empty stub; its .sln references a .csproj that doesn't exist in the repo.
  • MeetingIdMapper — one-off Python migration script backfilling rsMeetingId from legacy RS MySQL. Has hardcoded DB credentials in source — needs rotation/removal.
  • ProxyRelay — Docker container (squid/privoxy) fronting an authenticated upstream proxy; anti-bot IP-rotation infra. Not referenced by name in the scraper code — the scrapers' "Broker" component takes a generic proxy address, which may or may not point at a ProxyRelay instance (see Scraper & File Generator Process).
  • SilkGenerator — standalone batch job pre-warming silk images for recently-updated horses.
  • TroyenUploader — dev/ops utility for manually pushing test meeting/DataDump JSON into S3. Has a hardcoded S3 key/secret as a fallback default — needs rotation/removal.
  • EQUBasePDF — standalone PDF/HTML scraper for equibase.com (US racing), unrelated to the AU-focused pipeline.
  • CustomerDataTemplate.Tests — xUnit tests against Common/TroyenModels/TroyenAPI.
  • TroyenAPIBruno — Bruno API-client collection for manually exercising TroyenAPI during development.

Also flagged in TroyenRaceIngestor/docs/01-EXISTING-SYSTEM-ANALYSIS.md §4.4: Common/RsRaceCardHelper.cs has a hardcoded MySQL password as a fallback default.

TroyenRaceIngestor vs. the old scraper pipeline

Per TroyenRaceIngestor/docs/01-EXISTING-SYSTEM-ANALYSIS.md and 02-NEW-PROJECT-DESIGN.md:

  • Why it exists: the existing pipeline is three loosely-coupled processes (PunterScraper → S3 flag file → TroyenDataFilesGenerator, plus independent TroyenService cron jobs) with no orchestrator, no message queue, and no relational write path.
  • The pivot: instead of live scraping, TroyenRaceIngestor ingests pre-supplied "Meeting JSON" + "Race Data JSON" files (uploaded to S3 bucket punter-web-scraper) as an alternative input, aiming for output parity with PunterScraper's Mongo documents and RaceCard/DataDump files.
  • Strict independence by design: zero reference to PunterScraper or TroyenDataFilesGenerator; reuses only Common and SilkLibrary. It deliberately does not reference TroyenModels — it defines its own Meeting/Race/Horses/RaceCard/DataDump classes (matching [BsonElement] names for schema compatibility) and its own lock/frozen-field write-guard, since Common.MongoDbHelper's guard is gated on inheriting a TroyenModels base class.
  • It also reimplements race-card/comment/data-dump generation independently (own Generation/ module) and signals completion via a separate S3 bucket (troyen-gen), untouched from PunterScraper's troyon-watchdog mechanism.
  • Status: per the design doc this started as a phased implementation plan (scaffolding → Mongo write → race-card building → …). The project is wired into ProjectTroyon.sln with real source, but treat "how much of the design is actually implemented" as needing verification against current code rather than assumed complete. See [[project-troyenraceingestor]] context — check its mappers first if a field is missing for some meetings only, rather than modifying the older scrapers ([[feedback-dont-modify-old-system]]).

Deployment / ops

Two parallel deployment paths exist for the same "front" services, and it's unclear from the repo alone which is the actual production target (or whether they serve different environments):

  • Docker → GHCR: back-b&p.bat (TroyenDataFilesGenerator, PunterScraper, TroyenService), front-b&p.bat (TroyenAPI, TroyenDataHelpers, TroyenDataWeb), plus punter-scraper-b&p.bat / race-ingestor-b&p.bat. All push to ghcr.io/tpl-tech-titans/<name>:<version>. Compose files show long-lived Linux containers exposing ports 5100/5101/7501–7503.
  • IIS via GitHub Actions: .github/workflows/TD-DEV.yml (branch dev) and TD-PROD.yml (branch main) build/publish TroyenAPI → TroyenDataHelpers → TroyenDataWeb in sequence to Windows IIS on self-hosted runners.

Config/env (.env.sample): single Mongo DB (Troyon_Dev); self-hosted S3 at s3.troyendata.com (buckets punter-cache, troyen-resources, troyen-silk, punter-web-scraper, plus code-referenced troyon-watchdog/troyen-data/troyen-watcher/troyen-gen); internal/external/silk API base URLs; a scraper "Broker" relay (BROKER_URLS); headless-browser binaries; external scratchings/stats HTTP APIs; legacy RS MySQL credentials.

Starting points

  • Database Reference — MongoDB collection schema and the legacy RS MySQL tables/joins.
  • Scraper & File Generator Process — step-by-step deep dive on PunterScraper → S3 flag file → TroyenDataFilesGenerator, with a debugging checklist.
  • TroyenRaceIngestor/docs/01-EXISTING-SYSTEM-ANALYSIS.md — deep-dive on the current system's structure and gaps; the basis for the ingestor rewrite.
  • TroyenRaceIngestor/docs/02-NEW-PROJECT-DESIGN.md — design/implementation plan for the new ingestor.
  • PunterScraper/STANDALONE-APP.md — packaging notes for PunterScraperLauncher.
  • TroyenDataWeb/wwwroot/instruction-docs/ — end-user install guides for the standalone scraper tools.
  • CHANGELOG.neds-scraper.md, changelogs/ — per-service changelogs (Neds scraper, data-file generator, web, helpers).
  • User Guide — logging in, admin-portal navigation, and the scraper backup/fallback chain for data-entry staff.
Updated by Claude on Aug. 13, 2026, 4:18 a.m. · Task: updatewiki create DB documentation