Overview
Version 1 · Create initial project overview from repository README and layout
Troyen Beta — Racing Data Scrapers
Troyen Beta is a monorepo of racing data scrapers. Three independent .NET console apps each pull race data from a different source, transform it into a standardized DataDump (one per race) plus a meeting file (one per meeting), and upload both to S3 object storage, where a downstream intake worker consumes them.
- Repository: https://github.com/TPL-Tech-Titans/troyen-beta.git
- Solution:
troyen-beta.slnx(all .NET projects)
Scrapers
| Scraper | Source | Tech | Canonical bucket |
|---|---|---|---|
| PunterWebScraper | punters.com.au (website) | .NET 10 · Playwright (headed browser) | punter-web-scraper |
| NedsFormScraper | nedsform.com.au (website) | .NET 10 · HTTP (no browser) | neds-scraper |
| RsDataDump | RS database (MySQL) | .NET 8 · Dapper | ras-scraper |
All three also drop their generated files into the shared got bucket's pending/ folder for the downstream intake worker.
The common contract
Every scraper produces the same two standardized outputs:
DataDumpfile — one per race. The standardized race/runner record (meeting & race IDs, distance/weight conversions, jockey/trainer, performance statistics, past runs,rawStats).- Meeting file — one per meeting. The grouped intake format (meeting + per-race event details, entry conditions, weather, prize-money breakdown).
Both are written to two places: the canonical bucket for that source (keys mirroring {date}/{discipline}/{country}/{meeting}/…) and the got bucket's pending/ folder (flat, timestamped filenames) where the intake worker picks them up.
How it fits together
PunterWebScraper NedsFormScraper RsDataDump
punters.com.au nedsform.com.au RS database
│ │ │
└──── DataDump (per race) + meeting JSON ────┘
│
▼
S3 object storage (s3prod.troyendata.com)
punter-web-scraper neds-scraper ras-scraper
│
got ──► pending/ ──► intake worker
Projects
PunterWebScraper (.NET 10 · Playwright)
Scrapes meeting lists and race-detail pages for horses, greyhounds, and harness racing from punters.com.au, then generates DataDump + meeting files.
- Launches the browser headed (not headless) to avoid bot detection; locally reuses installed Microsoft Edge, falling back to Playwright's bundled Chromium on Linux/Docker.
- A VPN is recommended so requests originate from an expected network.
- "Skip if already scraped" checks the bucket, not local disk — identical behaviour locally and in a container.
dotnet run --project PunterWebScraper -- --date <yyyy-MM-dd> --discipline <horses|greyhounds|harness> [--country <name>]
--date and --discipline are required; --country narrows Phase 2 (race-detail scraping) to one country. Run with no arguments for interactive prompts (the mode the data-entry team uses).
NedsFormScraper (.NET 10 · HTTP, no browser)
nedsform.com.au is a Next.js app that server-renders its full data as an RSC payload, so the scraper reads structured JSON out of the page over plain HTTP — no browser required.
dotnet run -- list-meetings --date <yyyy-MM-dd>
dotnet run -- scrape-meeting --date <yyyy-MM-dd> --meeting "<venue>" --discipline <horse|harness|greyhound>
--meeting is fuzzy/case-insensitive. The RSC extraction is the most fragile part — see NedsFormScraper/CLAUDE.md.
RsDataDump (.NET 8 · Dapper)
Reads the RS MySQL database directly (no web scraping) and generates DataDump + meeting files for one meeting. Ships as rsdump.exe.
dotnet run --project RsDataDump -- --venue "<name|code>" --date <yyyy-MM-dd|today> --discipline <thoroughbred|harness|greyhound|arabian>
# diagnostics
dotnet run --project RsDataDump -- --check-s3 # list buckets the credentials can see
dotnet run --project RsDataDump -- --help
All three inputs are required. --venue accepts a fuzzy name or 4-letter code (e.g. Ipswich or IPSW). The data-entry team runs it via a run.bat prompt wrapper.
Storage (S3)
All scrapers write to S3-compatible object storage at s3prod.troyendata.com (path-style addressing). Each source has its own canonical bucket; all three also write to the shared got bucket's pending/ folder.
| Bucket | Written by | Contents |
|---|---|---|
punter-web-scraper |
PunterWebScraper | meeting list + per-race HTML/JSON, DataDump & meeting files |
neds-scraper |
NedsFormScraper | native meeting JSON + DataDump & meeting files |
ras-scraper |
RsDataDump | DataDump & meeting files |
got |
all three | pending/ DataDump + meeting files for the intake worker |
Note: PunterWebScraper's config section is historically named
Wasabi, but it points at thes3prod.troyendata.comS3 endpoint above — the same storage the other two use. It is not a Wasabi bucket.
Credentials
- Console builds shipped to the data-entry team have
appsettings.jsonembedded in the exe — no plaintext config ships. Values can be overridden per machine via an optionalappsettings.Local.jsonor environment variables. - Docker/scheduled runs supply credentials via environment variables / a secrets store (Rundeck Key Storage), never committed.
- Never commit real credentials. If any leak, rotate them and prefer least-privilege, bucket-scoped keys.
Building & running (developers)
Prerequisite: the .NET SDK 10.x builds all three console apps, including the net8.0 RsDataDump project.
dotnet build troyen-beta.slnx # build all projects
dotnet test troyen-beta.slnx # run all test projects
Distribution to the data-entry team
The three console scrapers are handed to non-technical operators as self-contained builds (the .NET runtime is bundled) that they run by double-clicking. Config/secrets are embedded in the exe, so no credential files travel in the download.
DEPLOYMENT.md— for developers: publish, package (zip), and share each build via Google Drive, plus the secrets model and a release checklist.INSTALL.md— for the data-entry team: download, extract, allow through SmartScreen, and run each tool (including the VPN requirement for PunterWebScraper).
Deployment (containers)
PunterWebScraper and NedsFormScraper each have a Dockerfile. The Punter scraper runs as a one-shot job:
docker build -t punter-web-scraper -f PunterWebScraper/Dockerfile .
docker run --rm --env-file PunterWebScraper/.env punter-web-scraper --date <yyyy-MM-dd> --discipline horses
Published image (GitHub Container Registry): ghcr.io/tpl-tech-titans/punter-web-scraper-beta. In production the Punter scraper runs on demand/schedule via Rundeck, with credentials pulled from Rundeck Key Storage. NedsFormScraper has its own Dockerfile + docker-compose.yml. RsDataDump is distributed as a desktop build rather than containerized.
Testing
- NedsFormScraper.Tests (xUnit) — RSC extraction, day-index & per-discipline race/runner parsing, timezone resolution, and the Troyon DataDump/meeting mapping, all against captured page fixtures (zero network).
- RsDataDump.Tests (xUnit) — the DataDump/meeting mapping from RS data.
dotnet test troyen-beta.slnx
Repository layout
troyen-beta/
├─ PunterWebScraper/ .NET 10 · Playwright scraper (punters.com.au) → punter-web-scraper
├─ NedsFormScraper/ .NET 10 · HTTP scraper (nedsform.com.au) → neds-scraper
│ └─ NedsFormScraper/ (the console app project; tests alongside)
├─ RsDataDump/ .NET 8 · MySQL → DataDump (RS database) → ras-scraper
├─ troyen-beta.slnx solution (all .NET projects)
├─ DEPLOYMENT.md developer release guide (build → zip → share)
├─ INSTALL.md data-entry team install & usage guide
└─ README.md repository entry point
Starting points
README.md— repository entry point (source of this overview).DEPLOYMENT.md— building, packaging, and sharing the console apps; secrets.INSTALL.md— end-user install & run guide for the data-entry team.NedsFormScraper/CLAUDE.md— deep technical reference for the Neds RSC extraction, Troyon mapping, and known fragile points.