← All projects
Software/Data Artifact untagged public artifact Updated: 2026-08-14

Knowledge Compiler

Structured evidence infrastructure with partial provenance recovery over a historical textbook-derived corpus.

The problem: Downstream systems need inspectable links between structured records and their source material, but textbooks are unstructured—buried in prose, tables, and figures.

The artifact: Knowledge Compiler implements a chunk → extract → normalize → structurally check pipeline. The historical corpus contains 2,305 files representing 2,294 unique IDs; 571 records meet the strongest recovered-subset criterion. Provenance recovery does not establish scientific validity or study admissibility.

Pipeline Architecture

Step 1
Chunk
Split textbook into digestible passages
Step 2
Extract
LLM extracts structured objects using ontology schema
Step 3
Normalize
Deduplicate and resolve cross-references
Step 4
Structural checks
Automated quality gates: schema and structural constraints
707
ACSM Objects
1,598
NSCA Objects
13
Ontology Types
5
Structural Checks

Ontology Structure

13 object types organized in a three-layer schema:

Core Concepts

  • • Concept
  • • Evidence
  • • Formula

Operational Rules

  • • Recommendation
  • • Threshold
  • • Procedure
  • • DecisionRule

Safety

  • • Warning
  • • Contraindication
  • • RiskFactor

Plus supporting types: Figure, Table, TableRow

FAQ

Why compile textbooks into structured knowledge?

Structured records can expose stored source locators to downstream systems. For the recovered subset, those locators support re-checkable source relations; they do not make the records ground truth or validated scientific evidence.

What's the ontology structure?

13 object types organized in a three-layer schema: Core Concepts (Concept, Evidence, Formula), Operational Rules (Recommendation, Threshold, Procedure, DecisionRule), and Safety (Warning, Contraindication, RiskFactor), plus supporting types (Figure, Table, TableRow).

Can this pipeline work for other textbooks?

The pipeline was applied to two exercise-science textbooks. That demonstrates implementation reuse within one domain; cross-domain generalization has not been tested.

Tech Stack

Python 3.11+ Pydantic LLM Extraction PyYAML OpenAI

Reproducibility

Env: Python 3.11+ with requirements.txt. Automated checks enforce schema and structural constraints; they do not establish factual accuracy, source fidelity, or scientific validity.

git clone https://github.com/MaxGuo1220/acsms12-manifest && cd acsms12-manifest && pip install -r requirements.txt && python -m compiler.cli run --book acsm12

Key Takeaways

  1. The historical inventory contains 2,305 files—707 from ACSM and 1,598 from NSCA—representing 2,294 unique IDs.
  2. 571 records meet the strongest recovered-subset criterion; 1,734 records are quarantined from unsupported reuse.
  3. Automated checks enforce structural constraints only; zero records currently have established scientific validity or study-admissible status.
  4. The pipeline was applied to two exercise-science textbooks, demonstrating transfer across sources within one domain.
  5. The pipeline is designed to be source-extensible; cross-domain generalization has not been tested.
  6. Pipeline cost: ~$0.003 per object using structured LLM extraction with schema enforcement.

Changelog

v2.02026-07-01 — NSCA CSCS 5th edition complete: 1,598 objects, 26 chapters.
v1.02026-05-01 — Historical ACSM extraction: 707 files.
v0.52026-03-01 — Pipeline v2: chunk → extract → normalize → validate with quality gates.
v0.12026-01-15 — Initial schema design and pilot extraction (15 objects).