Knowledge Compiler
Structured evidence infrastructure with partial provenance recovery over a historical textbook-derived corpus.
The problem: Downstream systems need inspectable links between structured records and their source material, but textbooks are unstructured—buried in prose, tables, and figures.
The artifact: Knowledge Compiler implements a chunk → extract → normalize → structurally check pipeline. The historical corpus contains 2,305 files representing 2,294 unique IDs; 571 records meet the strongest recovered-subset criterion. Provenance recovery does not establish scientific validity or study admissibility.
Pipeline Architecture
Ontology Structure
13 object types organized in a three-layer schema:
Core Concepts
- • Concept
- • Evidence
- • Formula
Operational Rules
- • Recommendation
- • Threshold
- • Procedure
- • DecisionRule
Safety
- • Warning
- • Contraindication
- • RiskFactor
Plus supporting types: Figure, Table, TableRow
FAQ
Why compile textbooks into structured knowledge?
Structured records can expose stored source locators to downstream systems. For the recovered subset, those locators support re-checkable source relations; they do not make the records ground truth or validated scientific evidence.
What's the ontology structure?
13 object types organized in a three-layer schema: Core Concepts (Concept, Evidence, Formula), Operational Rules (Recommendation, Threshold, Procedure, DecisionRule), and Safety (Warning, Contraindication, RiskFactor), plus supporting types (Figure, Table, TableRow).
Can this pipeline work for other textbooks?
The pipeline was applied to two exercise-science textbooks. That demonstrates implementation reuse within one domain; cross-domain generalization has not been tested.
Tech Stack
Reproducibility
Env: Python 3.11+ with requirements.txt. Automated checks enforce schema and structural constraints; they do not establish factual accuracy, source fidelity, or scientific validity.
git clone https://github.com/MaxGuo1220/acsms12-manifest && cd acsms12-manifest && pip install -r requirements.txt && python -m compiler.cli run --book acsm12 Key Takeaways
- The historical inventory contains 2,305 files—707 from ACSM and 1,598 from NSCA—representing 2,294 unique IDs.
- 571 records meet the strongest recovered-subset criterion; 1,734 records are quarantined from unsupported reuse.
- Automated checks enforce structural constraints only; zero records currently have established scientific validity or study-admissible status.
- The pipeline was applied to two exercise-science textbooks, demonstrating transfer across sources within one domain.
- The pipeline is designed to be source-extensible; cross-domain generalization has not been tested.
- Pipeline cost: ~$0.003 per object using structured LLM extraction with schema enforcement.