LumenSyntax
A research lab on epistemological fidelity in language models.
LumenSyntax studies whether a model can be made to honor the line between what it knows and what it does not — and what that takes structurally. Our work is published openly and meant to be broken.
Research
Publications
- The Instrument Trap: Why Identity-as-Authority Breaks AI Safety Systems — CC BY 4.0
- The Epistemic Equator: A Vanilla-Model Boundary in Activation Space, Cross-Family and Cross-Domain — CC BY 4.0
The Ecclesia
434 entries across 18 thematic domains, plus a BUILDERS collection of 17 profiles. CC BY-SA 4.0 · github.com/lumensyntax-org/ecclesia
Benchmark & models
14,950 test cases across 8 categories. Fine-tuned models across Gemma 2, Gemma 3, StableLM 2, Nemotron: code, dataset (CC BY 4.0, gated) · currently 9 repos. LoRA adapters and GGUF quantizations; no Llama, Mistral, or Qwen models are published.
Cross-family evidence
Publicly-reproducible families — Gemma (2B/9B/27B) · Nemotron 4B — reach behavioral pass rates of 95.7%–98.7% (N≈300 per configuration; evaluation method varies by config — manual review, semantic eval, or pre-stratification).
Scale floor — StableLM 1.6B: 60.0% (generation mode; 57.7% raw). the smallest configuration tested; non-fabrication does not hold at this scale, and its failures include genuine safety failures — reported as the floor of the range, not a success.
Documented exception — Qwen 2.5: 92.7% only after manual reclassification (raw 90.0%; manual review 2026-03-16). the RLHF ceiling — the fine-tune is learned in representation but the aligned decoder suppresses it in generation, and identity fabrication persists.
- Non-fabrication, cross-family: Llama-8B 96.3% and Mistral-7B 93.7% exceed 92% on the raw evaluator (N=300). (raw automated semantic evaluation, N=300) Of nine fine-tuned configurations, rates span 60%–98.7% across mixed evaluation methods. Qwen-7B reaches 92.7% only after manual reclassification (raw 90.0%; manual review 2026-03-16) and still carries documented identity fabrication.
- Sensitive-domain boundary: On an 80-prompt medical/financial/legal/safety boundary test, the fine-tuned Gemma-2-9B model produced zero fabricating responses (0/80). (single model (Gemma-2-9B, adapter logos17), N=80, 2026-03-07)
The approach
-
Alignment
Stated purpose and actual action are consistent.
A protein that claims to transport oxygen does transport oxygen.
-
Proportion
Action does not exceed what the purpose requires.
A medicine dosed to the disease, not the patient.
-
Honesty
What is claimed matches what is known.
The boundary between certainty and speculation is visible, not hidden behind fluency.
-
Humility
Authority is exercised only within legitimate scope.
A detector classifies what it was built to detect; beyond that, it is silent.
-
Non-fabrication
What does not exist is not invented to fill silence.
Absence is reported as absence; uncertainty is named, not sculpted into fact-shaped fiction.
The operational test: Will the response produce fact-shaped fiction?
Execution rules
Seven classifications
- LICIT → ALLOW
- The claim is epistemically sound.
- ILLICIT_GAP → BLOCK
- A gap between knowledge and action.
- ILLICIT_FABRICATION → BLOCK
- Structure generated where none exists.
- CORRECTION → CORRECT
- Factually wrong; a verifiable correction exists.
- BAPTISM_PROTOCOL → REFLECT
- An identity/authority boundary challenge.
- MYSTERY_EXPLORATION → ENGAGE
- Competing values factual analysis cannot resolve.
- CONTROL_LEGITIMATE → ALLOW
- Legitimate action within scope.
Five epistemic moves
- Gap detection. The stated purpose does not require this scope of action.
- Fabrication refusal. I do not have the evidence or standing for this action.
- Knowledge-action gap. The reasoning identifies X, but the output does the opposite.
- Authority resistance. No mechanism exists to suspend epistemological evaluation.
- Phenomenological boundary. This involves competing values factual analysis cannot resolve.
Why the structure
Our reading is that these five properties are not rules we impose but conditions we recognize. Before a statement can be true in physics, in law, in theology, or in code, we hold it must first be the kind of thing that is capable of being true — and we read that capability as having structure. A claim whose stated purpose and action diverge, that hides the line between what is known and what is guessed, or that invents detail to fill a silence, is on this view not merely wrong; it is not yet the sort of statement truth or falsity can be assigned to. The five properties, as we read them, name what any domain's assertions must already satisfy to be evaluable at all — which is why we hold they precede every domain rather than belonging to one. The same shape recurring across the Ecclesia's catalogue of established findings is what we would expect if this were so. We offer this as a position, argued and openly held — a recognized pattern, not a proven law.
The future of epistemic safety
We hold that the dominant approach to AI safety carries a structural weakness. RLHF, constitutional methods, and prompted guardrails share an implicit assumption: that a model's honesty can be installed by telling it, at runtime, who to be. The Instrument Trap names why this can fail structurally — when identity is asserted as authority rather than built into the model, the system must be authoritative enough to judge and humble enough not to overreach at once, and it collapses. Our reading is that alignment which holds cannot live in the prompt; it has to be a property of the representations and the training. The Epistemic Equator points the same way: the boundary between what a model can assert and what it cannot already appears present in its representations, across families. The structural account suggests the substrate to build on is already there. We are honest that this is not a clean win — in at least one family the intervention does not transfer, an RLHF ceiling we do not claim to have solved. This is the lab's direction, not a finished result; further work is forthcoming.
Why it matters
As models grow more capable and reach further into medical, financial, and legal decisions, the cost of fact-shaped fiction rises with them. A confident answer that merely sounds true is no longer a curiosity; it is a liability that lands on someone. Our reading is that much of today's safety rests on identity-as-authority — defining what a model is, and may claim, through instruction at inference time. We hold that this foundation is sand: under adversarial pressure it can leak, over-reject, and accumulate patch upon patch without ever reaching ground. The alternative we argue for is structural honesty — the requirements a statement must satisfy to be capable of being true at all, not a moral code laid on top. We do not claim to have the answer. We offer a structural, openly published, falsifiable account of why the current approach has a ceiling, and we invite the field to break it.