QDG Knowledge Base Read-only viewer QWebHub
notes

Status and Open Items

Version 1 · Publish the current state as at 7 September 2026 and the open work: reconcile not yet run, dedup review queue, scrape backlog, known source defects and the legacy staging decision.

Historical version

Status and Open Items

Last reviewed: 2026-09-07 — measured against the live databases.

Where things stand

The rebuild is done and shipped. rs.breeding holds 3,106,726 rows and qdb.breeding holds the same rows with identical ids. The table has been independently validated against four days of race-day runners with no error identified in it. See Overview for the full state table.

What has not happened: nothing has been locked or marked checked, and nothing downstream consumes a breeding id yet — qdb.runner.breeding_id is still 0 everywhere. That means the table is still provisional and can be rebuilt and re-shipped freely. The moment a runner carries a non-zero breeding_id, the ids become load-bearing and rs.breeding is effectively frozen. ship_to_qdb.py checks this guard rather than remembering it.


Open work

1. Reconcile has not been run against the rebuilt table

rebuild/reconcile_sources.py is written but not yet run. Until it does, is_locked and is_checked are 0 across all 3.1 M rows, and the roughly 387,165 dual-sourced rows are not identified as such — link_to_file records only the first source to claim a row.

This is the step that turns "both sources happen to agree" into recorded evidence. Run it before anything downstream starts depending on breeding ids.

2. Dedup review queue — 583 pairs

The September dedup merged 1,233 pairs and deliberately held back the rest for a person:

Verdict Pairs Progeny involved
REVIEW — duplicate has progeny, do not delete 537 1,888
REVIEW — foal date contradicts its own row 28 0
REVIEW — 3+ rows in consecutive years 18 0

The worklist is rs.breeding_dedup_20260905, which keeps each pair with its reason. The 537 with progeny are the ones that matter: merging them means repointing 1,888 real pedigree links, and a wrong one is invisible afterwards.

3. Anomalous links — 789 rows

reports/anomalous_links_sire.csv and _dam.csv. These are links where the parent's own year is more than one off the stated year, or any gap at all on a southern-hemisphere parent. They are name collisions rather than calendar differences — two horses sharing a name and a country — and rules R8–R11 cannot catch them because the age gaps stay plausible. This is the report to read first after any build.

4. Scrape backlog

Named but unlinked
Sires 4,819
Dams 54,611

Separately, 24,986 of 112,295 sires (22.2 %) have no scraped page of their own — including NORTHERN DANCER, NATIVE DANCER, SECRETARIAT, NASRULLAH, HYPERION and PHALARIS. They enter breeding correctly from the pages they appear on; what is lost is depth behind them. Dams are only 1.7 % missing.

reports/unlinked_sires.csv and unlinked_dams.csv rank this by how many children are waiting behind each missing horse, so the list is a work queue rather than an inventory.

5. Missing foals — backfill candidate

The August validation found five runners absent from both source feeds, roughly one horse in six hundred. Every parent of every missing foal was present and correct, and the race data holds the foaling date, sex, both parents and the breeder — so these records are reconstructable from the race data itself. Recommendation 4 of the validation report; not yet acted on.

6. Source defects needing a re-parse, not SQL

These cannot be fixed from the database. They need the source HTML re-read, which is a separate job on a separate schedule.

  • ~YYY- year truncation. PedigreeQuery's "approximate year, unknown final digit" notation is over-trimmed by one character, so ~184- parses as 184. year_text already holds the damaged value. This is why the table's minimum year_of_foal is 182, and it is the dominant cause of implausible parent-age gaps. A few hundred corrupt sire years block thousands of horses each — high-leverage to fix.
  • Turkish charset defect — Windows-1254 read as Latin-1.
  • ONMOUSE markup leak in publish_name. Caught by rule R5 at load, but the underlying parse is still wrong.

7. Legacy staging — 226 GB, keep or drop

Table Rows Size Remaining use
rs.parsed_horses 71,825,680 162 GB Input to step 1; a re-parse cache thereafter
rs.import_decisions 80,704,431 64 GB None
rs.breeding_candidate 501,596 0.2 GB Superseded staging table

import_decisions has no consumer at all now. The decision is Robert's; the disk is real.

8. Foreign keys

Not yet added. The loaders insert children before their parents exist, so FKs go on after a build — where the validating ALTER doubles as an integrity proof.

9. Documentation drift in the repository

HANDOVER.md (v2.0.0) and RUNBOOK_BREEDING_BUILD.md (v1.6.0) both date from 21 August and describe schema v3 with key_name, the eight-phase breeding_build.py command set, and an rs.breeding holding 88 proving rows. All three of those statements are now wrong. The README.md carries a "read first" note pointing at rebuild/, but the handover and runbook themselves have not been rewritten. This knowledge base is currently the more accurate description of the project.

Two documents referenced by the code are not in the repository at all: ANALYSIS_newAUData_dump.md (cited by load_studbook.py and dedup_adjacent_years.sql for the season-rule measurement) and DECISIONS.md (cited by archive/breedinginfill2/README.md, and the source of the C-series and F-series decision codes referenced in reconcile_sources.py). Worth locating and publishing here.


Resolved, for the record

  • The full build. It had never been run under the old pipeline; it now completes in about 31 minutes.
  • git push. Recorded as never having completed as of 21 August; the repository now carries six commits through b299932, including the 6 September consolidation.
  • The two-repository split. Closed 6 September; BreedinginFill2 retired with everything preserved under archive/breedinginfill2/.
  • Country codes on ODIN. qdb.country.racing_code corrected 17 August — 18 rows, 240 total, 67 codes, all distinct.
  • The 6,036 blocked rows of the old pipeline. Verified against source HTML; the flagged years are exactly what PedigreeQuery's own pages display. A source problem, not an import bug, and no defensible auto-fix exists. The rebuild replaced the mechanism that produced them.
Updated by Robert on Sept. 6, 2026, 11:46 p.m. · Commit: b299932