Product Evidence
Claims should survive contact with evidence.
Product evidence
This page presents product evidence for Phronelis Crucible.
The current public evidence case is a controlled full-cycle benchmark using BANKING77.
BANKING77 β Controlled Full-Cycle Benchmark
Challenge: a dataset can pass ordinary checks and still be unfit for training.
Crucible evaluated BANKING77 under a frozen benchmark contract covering exact within-split duplicates, exact train/test leakage, and exact classification contradictions.
The benchmark converted those checks into a reproducible technical decision for training handoff.
Workflow
Analyze -> Repair -> Owner Approval -> Curate -> Evidence Continuity -> Re-Audit A/B -> PASS -> Verified Training Handoff
- Repair actions
- 33
- Owner-approved
- 19
- Executed
- 19
- Original source
- Preserved
Result
- Critical leakage groups: 7 -> 0.
- Train/eval leakage reduced to 0.
- Active duplicate groups: 12 -> 0.
- Contradiction groups: 0 -> 0.
- TrainingGap: 0.470725 -> 0.070438.
- TrainingReadiness: 0.0 -> 0.929562.
- UsefulInformationScore: 0.529275 -> 0.929562.
Verification and integrity
- Re-audit A: PASS.
- Re-audit B: PASS.
- Substantive equivalence: Confirmed.
- Failed hard gates: 0.
- Final verdict: PASS.
- Training handoff: READY_FOR_TRAINING_HANDOFF.
- Independent rebuilds were byte-identical.
- Manifest and checksum coverage: PASS.
- Package verification: PASS.
Conclusion
Under the controlled BANKING77 benchmark, Crucible completed the full audit-to-curation-to-verification workflow reproducibly and produced a verified training handoff.
Benchmark scope: exact within-split duplicates, exact train/test leakage and exact classification contradictions under the frozen BANKING77 benchmark contract.