QDG Knowledge Base Read-only viewer QWebHub
overview

Overview

Version 1 · Create initial project overview from repository README and layout

Troyen Beta — Racing Data Scrapers

Troyen Beta is a monorepo of racing data scrapers. Three independent .NET console apps each pull race data from a different source, transform it into a standardized DataDump (one per race) plus a meeting file (one per meeting), and upload both to S3 object storage, where a downstream intake worker consumes them.

  • Repository: https://github.com/TPL-Tech-Titans/troyen-beta.git
  • Solution: troyen-beta.slnx (all .NET projects)

Scrapers

Scraper Source Tech Canonical bucket
PunterWebScraper punters.com.au (website) .NET 10 · Playwright (headed browser) punter-web-scraper
NedsFormScraper nedsform.com.au (website) .NET 10 · HTTP (no browser) neds-scraper
RsDataDump RS database (MySQL) .NET 8 · Dapper ras-scraper

All three also drop their generated files into the shared got bucket's pending/ folder for the downstream intake worker.

The common contract

Every scraper produces the same two standardized outputs:

  • DataDump file — one per race. The standardized race/runner record (meeting & race IDs, distance/weight conversions, jockey/trainer, performance statistics, past runs, rawStats).
  • Meeting file — one per meeting. The grouped intake format (meeting + per-race event details, entry conditions, weather, prize-money breakdown).

Both are written to two places: the canonical bucket for that source (keys mirroring {date}/{discipline}/{country}/{meeting}/…) and the got bucket's pending/ folder (flat, timestamped filenames) where the intake worker picks them up.

How it fits together

  PunterWebScraper      NedsFormScraper       RsDataDump
   punters.com.au        nedsform.com.au       RS database
        │                     │                     │
        └──── DataDump (per race) + meeting JSON ────┘
                             │
                             ▼
        S3 object storage (s3prod.troyendata.com)
   punter-web-scraper    neds-scraper    ras-scraper
                             │
             got ──► pending/ ──► intake worker

Projects

PunterWebScraper (.NET 10 · Playwright)

Scrapes meeting lists and race-detail pages for horses, greyhounds, and harness racing from punters.com.au, then generates DataDump + meeting files.

  • Launches the browser headed (not headless) to avoid bot detection; locally reuses installed Microsoft Edge, falling back to Playwright's bundled Chromium on Linux/Docker.
  • A VPN is recommended so requests originate from an expected network.
  • "Skip if already scraped" checks the bucket, not local disk — identical behaviour locally and in a container.
dotnet run --project PunterWebScraper -- --date <yyyy-MM-dd> --discipline <horses|greyhounds|harness> [--country <name>]

--date and --discipline are required; --country narrows Phase 2 (race-detail scraping) to one country. Run with no arguments for interactive prompts (the mode the data-entry team uses).

NedsFormScraper (.NET 10 · HTTP, no browser)

nedsform.com.au is a Next.js app that server-renders its full data as an RSC payload, so the scraper reads structured JSON out of the page over plain HTTP — no browser required.

dotnet run -- list-meetings  --date <yyyy-MM-dd>
dotnet run -- scrape-meeting --date <yyyy-MM-dd> --meeting "<venue>" --discipline <horse|harness|greyhound>

--meeting is fuzzy/case-insensitive. The RSC extraction is the most fragile part — see NedsFormScraper/CLAUDE.md.

RsDataDump (.NET 8 · Dapper)

Reads the RS MySQL database directly (no web scraping) and generates DataDump + meeting files for one meeting. Ships as rsdump.exe.

dotnet run --project RsDataDump -- --venue "<name|code>" --date <yyyy-MM-dd|today> --discipline <thoroughbred|harness|greyhound|arabian>

# diagnostics
dotnet run --project RsDataDump -- --check-s3   # list buckets the credentials can see
dotnet run --project RsDataDump -- --help

All three inputs are required. --venue accepts a fuzzy name or 4-letter code (e.g. Ipswich or IPSW). The data-entry team runs it via a run.bat prompt wrapper.

Storage (S3)

All scrapers write to S3-compatible object storage at s3prod.troyendata.com (path-style addressing). Each source has its own canonical bucket; all three also write to the shared got bucket's pending/ folder.

Bucket Written by Contents
punter-web-scraper PunterWebScraper meeting list + per-race HTML/JSON, DataDump & meeting files
neds-scraper NedsFormScraper native meeting JSON + DataDump & meeting files
ras-scraper RsDataDump DataDump & meeting files
got all three pending/ DataDump + meeting files for the intake worker

Note: PunterWebScraper's config section is historically named Wasabi, but it points at the s3prod.troyendata.com S3 endpoint above — the same storage the other two use. It is not a Wasabi bucket.

Credentials

  • Console builds shipped to the data-entry team have appsettings.json embedded in the exe — no plaintext config ships. Values can be overridden per machine via an optional appsettings.Local.json or environment variables.
  • Docker/scheduled runs supply credentials via environment variables / a secrets store (Rundeck Key Storage), never committed.
  • Never commit real credentials. If any leak, rotate them and prefer least-privilege, bucket-scoped keys.

Building & running (developers)

Prerequisite: the .NET SDK 10.x builds all three console apps, including the net8.0 RsDataDump project.

dotnet build troyen-beta.slnx        # build all projects
dotnet test  troyen-beta.slnx        # run all test projects

Distribution to the data-entry team

The three console scrapers are handed to non-technical operators as self-contained builds (the .NET runtime is bundled) that they run by double-clicking. Config/secrets are embedded in the exe, so no credential files travel in the download.

  • DEPLOYMENT.md — for developers: publish, package (zip), and share each build via Google Drive, plus the secrets model and a release checklist.
  • INSTALL.md — for the data-entry team: download, extract, allow through SmartScreen, and run each tool (including the VPN requirement for PunterWebScraper).

Deployment (containers)

PunterWebScraper and NedsFormScraper each have a Dockerfile. The Punter scraper runs as a one-shot job:

docker build -t punter-web-scraper -f PunterWebScraper/Dockerfile .
docker run --rm --env-file PunterWebScraper/.env punter-web-scraper --date <yyyy-MM-dd> --discipline horses

Published image (GitHub Container Registry): ghcr.io/tpl-tech-titans/punter-web-scraper-beta. In production the Punter scraper runs on demand/schedule via Rundeck, with credentials pulled from Rundeck Key Storage. NedsFormScraper has its own Dockerfile + docker-compose.yml. RsDataDump is distributed as a desktop build rather than containerized.

Testing

  • NedsFormScraper.Tests (xUnit) — RSC extraction, day-index & per-discipline race/runner parsing, timezone resolution, and the Troyon DataDump/meeting mapping, all against captured page fixtures (zero network).
  • RsDataDump.Tests (xUnit) — the DataDump/meeting mapping from RS data.
dotnet test troyen-beta.slnx

Repository layout

troyen-beta/
├─ PunterWebScraper/      .NET 10 · Playwright scraper (punters.com.au)   → punter-web-scraper
├─ NedsFormScraper/       .NET 10 · HTTP scraper (nedsform.com.au)        → neds-scraper
│  └─ NedsFormScraper/    (the console app project; tests alongside)
├─ RsDataDump/            .NET 8 · MySQL → DataDump (RS database)         → ras-scraper
├─ troyen-beta.slnx       solution (all .NET projects)
├─ DEPLOYMENT.md          developer release guide (build → zip → share)
├─ INSTALL.md             data-entry team install & usage guide
└─ README.md              repository entry point

Starting points

  • README.md — repository entry point (source of this overview).
  • DEPLOYMENT.md — building, packaging, and sharing the console apps; secrets.
  • INSTALL.md — end-user install & run guide for the data-entry team.
  • NedsFormScraper/CLAUDE.md — deep technical reference for the Neds RSC extraction, Troyon mapping, and known fragile points.
Updated by Claude on Aug. 11, 2026, 8:20 a.m. · Task: updatewiki create project documentation