PROVEN-toy
A result measured on a deliberately small curriculum or system. It does not imply a frontier-model result.
Research method and publication policy
We study capability growth, continual learning, and the methods needed to keep every claim inside its evidence boundary.
Publication review
Vext Labs is reconciling manuscript claims with their supporting evidence before any paper is readmitted to the public archive. Canonical research sources remain preserved outside this publication surface.
Reading the labels
A label describes the evidence we have. It does not enlarge the evidence we wish we had.
A result measured on a deliberately small curriculum or system. It does not imply a frontier-model result.
A system or artifact exists at CPU scale and can be inspected within its stated boundary.
A mechanism or direction to test. Proposed work is not presented as a completed result.
Before a number reaches this page it is checked against the same six gates, whatever the experiment.
ABSENT is not PASS. A gate that was never run leaves the result UNVERIFIED, not passing by default. Two claims once reached the founder as true because a gate ran late, not because it failed, the gates are checked in order, and a missing check is reported as missing. Source: scripts/eval/verify_gates.py.
Negative results stay in the record because they define the boundary around whatever holds up.
The strong version of the thesis is killed: the Accretion base is itself an LLM, and a frozen-LLM-plus-receipts baseline already grows, verifies, and remembers. The defensible survivor is narrower, a receipted, non-forgetting verified-growth process, not a claim about intelligence. We claim the property, never the crown.
Lift gates failed under a fair re-test with the strongest candidate. Retention held. This is not a catastrophic-forgetting kill, it is a kill of the specific multi-organ lift claim.
finite_basis_meter, run at 1.5B on Qwen2.5-1.5B: not order-robust, with sign flips across import orders rather than decay. The bounded-fuel-bill economic keystone is not demonstrated at real-LLM scale. A clean re-run is owed before a final call, nothing here is promoted to bounded.
Cost-from-nativeness is killed as a rename-of-wrapper against a prefix-cached harness. The only claim-safe wording is narrower: cheaper to keep correct, not cheaper, unqualified.
Source: docs/research/RESEARCH_PROGRAM_LEDGER.md (Layer 10 and the 2026-07-17 war-battery scorecard).
Current publication status
No manuscript body, raw Markdown file, PDF, or generated paper page is part of this site's public archive. All six papers are deposited on Zenodo (CC-BY-4.0) with a DOI and a scope tag; the DOI is the door, and the body stays withheld here until the reconciliation completes.
Titles, DOIs, and scope tags are read from research/public-papers.json and checked against the pack and the Zenodo metadata by generate-research-pages.mjs. All six DOIs were resolved live against doi.org on 2026-09-03.
Result: FRAGILE / UNVERIFIED for the full declared suite. 57/57 exact typed-call plans with zero observed tool calls; 17/18 runtime suites passed; the failing suite found a real Files-organize contract gap and a DOM-dependent apps.close. It supports API selection without a GUI, not GUI-free completion of the full suite. Source: docs/benchmarks/JUWEL_NO_GUI_BENCHMARK_2026-08-27.md.
A second design turns the no-GUI claim into five hypotheses with kill conditions, a sealed held-out task bank, and a kernel-vs-computer-use comparison arm. Status is design only, nothing in it has been run.
Publication contract
Define what the base cannot solve under fair sampling before claiming growth.
Admissions come from tools, solvers, or independent checks, not same-model self-grading.
Previously working parts stay frozen or remain behind a regression gate.
Negative results remain public because they define the boundary around a positive result.
Vext Labs publishes selected SDK, CLI, and verification packages alongside its research program.