Overview
Version 2 · Correct TroyenDataWeb auth description: login is custom email+TOTP against MongoDB (AuthService/TotpService), not ASP.NET Identity/EF Core — confirmed by reading LoginController/AccountController/UserManagementController
Historical versionTroyon Data — Racing Data Platform
Troyon Data ("Troyen"/"Troyon" are used interchangeably across the codebase — e.g. TroyenService builds as TroyonService) is a thoroughbred / harness / greyhound racing data pipeline and content platform.
- Repository: https://github.com/TPL-Tech-Titans/Troyon-Data.git
- Solution:
ProjectTroyon.sln(only 12 of ~25 top-level folders are wired into it — the rest are standalone tools/scripts/legacy stubs) - .NET SDK: pinned to
8.0.0viaglobal.json(rollForward: latestMajor)
What it does
The system scrapes race-meeting data from bookmaker/aggregator sites — punters.com.au + racenet (PunterScraper), neds.com.au (NedsScraper) — and cross-references it against a legacy read-only MySQL system called "RS" (Racing & Sports). Everything is persisted in MongoDB (Troyon_Dev DB) as meetings / races (embedding runners) / horses documents.
A second stage (TroyenDataFilesGenerator) watches S3 for "flag files" dropped by PunterScraper and turns raw meeting/race data into polished, LLM-enriched Race Cards and Data Dumps (JSON) — including AI-generated runner comments/spotlights (via LlmTornado/OpenAI, 3 language variants) and generated jockey silk colour images (SVG→PNG via SilkLibrary). Client-specific ("RS-client") race-card variants are also produced by reading the legacy RS MySQL directly.
Results are served externally via TroyenAPI (read-only; auth is a custom API-key scheme, not JWT — see TroyenAPI Reference) and internally via TroyenDataHelpers (read/write admin API, no enforced auth in code — see TroyenDataHelpers Reference), with TroyenDataWeb acting as the admin/ops portal for managing meetings/races/silks/comments. TroyenService runs three independent scheduled jobs (Windows Service, Cronos-scheduled) that enrich the same Mongo documents: RS-ID cross-reference, runner stats, scratchings.
Data flow
PunterScraper ──(S3 flag file: troyon-watchdog)──► TroyenDataFilesGenerator
(punters.com.au + (builds RaceCard/DataDump JSON,
racenet, → MongoDB) LLM comments, silks, RS-client cards)
│
NedsScraper ──(independent, own Web API)──► MongoDB ◄────────┘
(neds.com.au) │
▼
TroyenAPI (external, read-only) ◄─┐
TroyenDataHelpers (internal, r/w) │
│ │
▼ │
TroyenDataWeb (admin portal, calls both)
TroyenService (independent cron jobs: RS-ID sync, runner stats, scratchings) ──► MongoDB
PunterScraper and TroyenDataFilesGenerator communicate only via the S3 flag file — there is no orchestrator or message queue, and MongoDB is the only write-side datastore in the core pipeline (RS MySQL is legacy and read-only).
Core projects (in ProjectTroyon.sln)
| Project | Responsibility | Tech |
|---|---|---|
| Common | Shared foundation: MongoDB access, S3 storage, LLM content generation, RS-MySQL race-card helper, scraper HTTP/broker helpers. Referenced by nearly everything. | net8.0 lib — AWSSDK.S3, MongoDB.Driver, MySqlConnector, Selenium, LlmTornado |
| TroyenModels | Shared BSON domain models (Meeting, Race, Horses, RaceCard, DataDump) used across the Mongo-backed pipeline. |
net8.0 lib — EF Core, MongoDB.Driver |
| TroyenData.Core | Racing.PaceCalculator — effectively dead code, only referenced from commented-out scraper code. |
net8.0 lib, no packages |
| PunterScraper | Production scraper: one-shot console/Docker job hitting punters.com.au GraphQL for meetings/races/horses/lookups, writes to Mongo, drops an S3 flag file. | net8.0 console — refs Common, TroyenModels, TroyenData.Core. Docker: punter-scraper |
| NedsScraper | Actively-maintained Web API scraper for neds.com.au; own changelog shows weekly development. Independent of the PunterScraper flow, writes the same Mongo collections. | net8.0 Web API — HtmlAgilityPack, LlmTornado, Playwright, Serilog |
| SilkLibrary | Renders jockey/horse "silk" colour SVGs to PNG (SkiaSharp), uploads to S3. | net8.0 lib — SkiaSharp, Svg.Skia |
| TroyenDataFilesGenerator | The "race card factory": polls S3 for PunterScraper flag files, builds RaceCard/DataDump JSON per discipline, runs LLM comment/spotlight generation, generates silks, builds RS-client variants. |
net8.0 Web/Exe. Docker: troyen-data-gen |
| TroyenAPI | External, read-only customer-facing REST API for meetings/races/racecards. Custom API-key auth (not JWT), API versioning. IIS site TroyonAPI/TD-External. |
net8.0 Web, win-x64 self-contained. Docker: troyen-external-api |
| TroyenDataHelpers | Internal read/write admin API (Meetings/RaceUpdate/Races/Horses controllers) — backend behind TroyenDataWeb. Bundles silk pattern SVGs. No enforced auth in code. IIS TroyenDataHelpersAPI/TD-Internal. |
net8.0 Web, win-x64. Docker: troyen-internal-api |
| TroyenDataWeb | Internal Admin Portal: ASP.NET MVC, admin views for runners/races/comments/silks. Login is custom email + TOTP against MongoDB (AuthService/TotpService, not ASP.NET Identity/EF Core, despite EF Core/Pomelo.MySql packages being referenced in the project). Hosts install-guide docs for the standalone scraper tools. IIS TroyenDataWeb/TD-Web. |
net8.0 Web, win-x64. Docker: troyen-web |
| TroyenService | Windows Service running 3 independent cron jobs against Mongo: RS-ID sync, runner stats, scratchings. Not an orchestrator — no reference to PunterScraper or TroyenDataFilesGenerator. | net8.0 Worker — Cronos. Docker: troyen-services |
| TroyenRaceIngestor | New, parallel ingestion pipeline (see below). | net8.0 Web (embedded Kestrel queue API). Docker: troyen-race-ingestor, port 5101 |
Standalone tools / legacy (not in .sln)
- PunterScraperLauncher — WPF Windows GUI wrapping a self-contained
PunterScraper.exebuild for non-technical staff; must be manually rebundled whenPunterScraper/Commonchanges. - PunterWebScraper — DOM-based (HtmlAgilityPack) punters.com.au form-guide scraper, distributed as a VPN-gated zip; a plausible successor to a Selenium prototype but not wired into the main pipeline yet.
- RacenetScraper — stale (~1 year); current
Program.csis a small Mongo batch query, not an active racenet.com.au scraper. - RacingPostscraper — empty stub; its
.slnreferences a.csprojthat doesn't exist in the repo. - MeetingIdMapper — one-off Python migration script backfilling
rsMeetingIdfrom legacy RS MySQL. Has hardcoded DB credentials in source — needs rotation/removal. - ProxyRelay — Docker container (squid/privoxy) fronting an authenticated upstream proxy; anti-bot IP-rotation infra supporting the scrapers.
- SilkGenerator — standalone batch job pre-warming silk images for recently-updated horses.
- TroyenUploader — dev/ops utility for manually pushing test meeting/DataDump JSON into S3. Has a hardcoded S3 key/secret as a fallback default — needs rotation/removal.
- EQUBasePDF — standalone PDF/HTML scraper for equibase.com (US racing), unrelated to the AU-focused pipeline.
- CustomerDataTemplate.Tests — xUnit tests against Common/TroyenModels/TroyenAPI.
- TroyenAPIBruno — Bruno API-client collection for manually exercising
TroyenAPIduring development.
Also flagged in
TroyenRaceIngestor/docs/01-EXISTING-SYSTEM-ANALYSIS.md§4.4:Common/RsRaceCardHelper.cshas a hardcoded MySQL password as a fallback default.
TroyenRaceIngestor vs. the old scraper pipeline
Per TroyenRaceIngestor/docs/01-EXISTING-SYSTEM-ANALYSIS.md and 02-NEW-PROJECT-DESIGN.md:
- Why it exists: the existing pipeline is three loosely-coupled processes (
PunterScraper→ S3 flag file →TroyenDataFilesGenerator, plus independentTroyenServicecron jobs) with no orchestrator, no message queue, and no relational write path. - The pivot: instead of live scraping,
TroyenRaceIngestoringests pre-supplied "Meeting JSON" + "Race Data JSON" files (uploaded to S3 bucketpunter-web-scraper) as an alternative input, aiming for output parity withPunterScraper's Mongo documents and RaceCard/DataDump files. - Strict independence by design: zero reference to
PunterScraperorTroyenDataFilesGenerator; reuses onlyCommonandSilkLibrary. It deliberately does not referenceTroyenModels— it defines its ownMeeting/Race/Horses/RaceCard/DataDumpclasses (matching[BsonElement]names for schema compatibility) and its own lock/frozen-field write-guard, sinceCommon.MongoDbHelper's guard is gated on inheriting aTroyenModelsbase class. - It also reimplements race-card/comment/data-dump generation independently (own
Generation/module) and signals completion via a separate S3 bucket (troyen-gen), untouched fromPunterScraper'stroyon-watchdogmechanism. - Status: per the design doc this started as a phased implementation plan (scaffolding → Mongo write → race-card building → …). The project is wired into
ProjectTroyon.slnwith real source, but treat "how much of the design is actually implemented" as needing verification against current code rather than assumed complete. See [[project-troyenraceingestor]] context — check its mappers first if a field is missing for some meetings only, rather than modifying the older scrapers ([[feedback-dont-modify-old-system]]).
Deployment / ops
Two parallel deployment paths exist for the same "front" services, and it's unclear from the repo alone which is the actual production target (or whether they serve different environments):
- Docker → GHCR:
back-b&p.bat(TroyenDataFilesGenerator,PunterScraper,TroyenService),front-b&p.bat(TroyenAPI,TroyenDataHelpers,TroyenDataWeb), pluspunter-scraper-b&p.bat/race-ingestor-b&p.bat. All push toghcr.io/tpl-tech-titans/<name>:<version>. Compose files show long-lived Linux containers exposing ports 5100/5101/7501–7503. - IIS via GitHub Actions:
.github/workflows/TD-DEV.yml(branchdev) andTD-PROD.yml(branchmain) build/publishTroyenAPI→TroyenDataHelpers→TroyenDataWebin sequence to Windows IIS on self-hosted runners.
Config/env (.env.sample): single Mongo DB (Troyon_Dev); self-hosted S3 at s3.troyendata.com (buckets punter-cache, troyen-resources, troyen-silk, punter-web-scraper, plus code-referenced troyon-watchdog/troyen-data/troyen-watcher/troyen-gen); internal/external/silk API base URLs; a scraper "Broker" relay (BROKER_URLS); headless-browser binaries; external scratchings/stats HTTP APIs; legacy RS MySQL credentials.
Starting points
TroyenRaceIngestor/docs/01-EXISTING-SYSTEM-ANALYSIS.md— deep-dive on the current system's structure and gaps; the basis for the ingestor rewrite.TroyenRaceIngestor/docs/02-NEW-PROJECT-DESIGN.md— design/implementation plan for the new ingestor.PunterScraper/STANDALONE-APP.md— packaging notes forPunterScraperLauncher.TroyenDataWeb/wwwroot/instruction-docs/— end-user install guides for the standalone scraper tools.CHANGELOG.neds-scraper.md,changelogs/— per-service changelogs (Neds scraper, data-file generator, web, helpers).- User Guide — logging in, admin-portal navigation, and the scraper backup/fallback chain for data-entry staff.