Skip to content
Vext Labs publishes "Beyond the Scaling Ceiling." Read the paper
Vext Labs

Research method and publication policy

Evidence beside every claim.

We study capability growth, continual learning, and the methods needed to keep every claim inside its evidence boundary.

Publication review

Public manuscripts are temporarily withheld.

Vext Labs is reconciling manuscript claims with their supporting evidence before any paper is readmitted to the public archive. Canonical research sources remain preserved outside this publication surface.

  • PROVEN
  • BUILT
  • REGISTRY-ONLY
  • GPU-gated
  • NOT-BUILT
  • UNAVAILABLE
  • KILLED

Reading the labels

Scope is part of the result.

A label describes the evidence we have. It does not enlarge the evidence we wish we had.

Measured

PROVEN-toy

A result measured on a deliberately small curriculum or system. It does not imply a frontier-model result.

Constructed

BUILT-CPU

A system or artifact exists at CPU scale and can be inspected within its stated boundary.

Open question

Proposed

A mechanism or direction to test. Proposed work is not presented as a completed result.

KILLEDFalsified on the bench. An honest kill stays in the record, not a rewrite.
OPENDesigned and pre-registered. Not yet run.
GPU-gatedNeeds a real LLM run. The bridge from toy to real is priced, not crossed.
BANKEDA prior result this program builds on without re-deriving.

Six gates, run against every result.

Before a number reaches this page it is checked against the same six gates, whatever the experiment.

  • G1

    Null control

  • G2

    Positive control

  • G3

    Variance

  • G4

    Exclusions

  • G5

    Prior art

  • G6

    Measurement

ABSENT is not PASS. A gate that was never run leaves the result UNVERIFIED, not passing by default. Two claims once reached the founder as true because a gate ran late, not because it failed, the gates are checked in order, and a missing check is reported as missing. Source: scripts/eval/verify_gates.py.

A kill narrows the search. It does not get quietly dropped.

Negative results stay in the record because they define the boundary around whatever holds up.

  • KILLED

    "LLMs are not intelligence; the Accretion Model is."

    The strong version of the thesis is killed: the Accretion base is itself an LLM, and a frozen-LLM-plus-receipts baseline already grows, verifies, and remembers. The defensible survivor is narrower, a receipted, non-forgetting verified-growth process, not a claim about intelligence. We claim the property, never the crown.

  • KILL

    Multiorgan / multi-faculty entity

    Lift gates failed under a fair re-test with the strongest candidate. Retention held. This is not a catastrophic-forgetting kill, it is a kill of the specific multi-organ lift claim.

  • KILL-UNBOUNDED

    Bounded fuel bill, first real-LLM rung

    finite_basis_meter, run at 1.5B on Qwen2.5-1.5B: not order-robust, with sign flips across import orders rather than decay. The bounded-fuel-bill economic keystone is not demonstrated at real-LLM scale. A clean re-run is owed before a final call, nothing here is promoted to bounded.

  • KILLED

    "Cheaper," unqualified

    Cost-from-nativeness is killed as a rename-of-wrapper against a prefix-cached harness. The only claim-safe wording is narrower: cheaper to keep correct, not cheaper, unqualified.

Source: docs/research/RESEARCH_PROGRAM_LEDGER.md (Layer 10 and the 2026-07-17 war-battery scorecard).

Current publication status

No manuscript body is public. The citation is.

No manuscript body, raw Markdown file, PDF, or generated paper page is part of this site's public archive. All six papers are deposited on Zenodo (CC-BY-4.0) with a DOI and a scope tag; the DOI is the door, and the body stays withheld here until the reconciliation completes.

  • VL-2026-01 Beyond the Scaling Ceiling PROVEN-TOY
    DOI 10.5281/zenodo.21628058 workshop draftPROVEN-toyBUILT/VERIFIED-CPUnot real-LLM scale for CIP forgetting product claims Body withheld on this site pending evidence reconciliation.
  • VL-2026-02 The Accretion Model BUILT/VERIFIED-CPU
    DOI 10.5281/zenodo.21628429 BUILT/VERIFIED-CPUPROVEN-toysystems and trust, not a capability crown Body withheld on this site pending evidence reconciliation.
  • VL-2026-03 The Leakage Signature REPRODUCIBLE ARTIFACT
    DOI 10.5281/zenodo.21628524 reproducible artifacttoy + bounded real-LLM addendum (1.5B to 32B demos)no frontier parity claim Body withheld on this site pending evidence reconciliation.
  • VL-2026-04 Verifier-Centric Capability Growth MEASURED + PROPOSED
    DOI 10.5281/zenodo.21628544 Part IV is a proposed research program and reports no results Body withheld on this site pending evidence reconciliation.
  • VL-2026-05 Sound Compounding Ratchet PROVEN-TOY
    DOI 10.5281/zenodo.21628552 all positive results are CPU-scale kill-teststhe real-LLM experiment is pre-registered and unrun Body withheld on this site pending evidence reconciliation.
  • VL-2026-06 Capability Injection via Reverse Abliteration METHOD PAPER
    DOI 10.5281/zenodo.21628566 method paperno production deployment claim is made Body withheld on this site pending evidence reconciliation.

Titles, DOIs, and scope tags are read from research/public-papers.json and checked against the pack and the Zenodo metadata by generate-research-pages.mjs. All six DOIs were resolved live against doi.org on 2026-09-03.

  • JUWEL OS zero-GUI agent benchmark FRAGILE

    Result: FRAGILE / UNVERIFIED for the full declared suite. 57/57 exact typed-call plans with zero observed tool calls; 17/18 runtime suites passed; the failing suite found a real Files-organize contract gap and a DOM-dependent apps.close. It supports API selection without a GUI, not GUI-free completion of the full suite. Source: docs/benchmarks/JUWEL_NO_GUI_BENCHMARK_2026-08-27.md.

  • JUWEL OS benchmark v2, a falsifiable design DESIGN · NOT RUN

    A second design turns the no-GUI claim into five hypotheses with kill conditions, a sealed held-out task bank, and a kernel-vs-computer-use comparison arm. Status is design only, nothing in it has been run.

Publication contract

The method travels with the claim.

State

Name the wall.

Define what the base cannot solve under fair sampling before claiming growth.

Check

Use a distinct verifier.

Admissions come from tools, solvers, or independent checks, not same-model self-grading.

Preserve

Guard prior work.

Previously working parts stay frozen or remain behind a regression gate.

Record

Keep the kill.

Negative results remain public because they define the boundary around a positive result.

Inspect the tools behind the record.

Vext Labs publishes selected SDK, CLI, and verification packages alongside its research program.