Sneferu.ai, Inc. · August 2026

The operating system for finished work.

Autonomous work that ships with receipts.

Founder-built. Local-first. Multi-model. Building new things right now.
The Sneferu console, live: a research run executing, its event log, scores, and the searchable run archive
LiveThe console, August 18, 2026: a research run, seven hours in, hunting for an exoplanet: every model call, every check, and every score written down as it happens.
01The premise

Generation is easy. Finishing is not.

An AI model should never be the sole judge of its own output.

A persuasive response is not evidence. Agreement from the same underlying intelligence is not independent review. Completion must be earned through challenge, mechanical checks, and preserved evidence.

When the work is done, Sneferu delivers it. When it is not, it says what remains.

AI can generate answers. Sneferu makes the answer survive an argument.
The argument, live: five critics attack each draft; two graders disagree; a running list tracks the problems the work must fix
The argument, liveFive different critics attack each draft, turn after turn; two graders score it and their disagreement is shown rather than averaged away; on the right, the running list of problems the work had to fix before it was allowed to finish.
02The system

Tell Sneferu what to build. The work comes back done.

Every workflow is a production line you start the same way: describe what you want, set a budget, and let it run.

The system map: every production line and gate in one diagram
The whole system on one pageEvery line below, drawn as it actually runs. You are seeing 47 of 809 nodes.
Open the map and click around →
Business Foundry

A rough idea to a business with a buyer in mind.

Twenty stages run without anyone steering: find the opportunity, kill the weaker options in a tournament, write the commercial case, build the product, and pack the launch.

Life Skills Foundations and Restoration Copilot came off this line.
Research Expedition

A question to a finding the data supports.

Testable directions, each grounded in the literature, then real experiments on real public data. The numbers decide, and what they will not support is thrown out.

The Z boson at 201σ from raw CERN data, five known planets recovered from telescope archives, and one 5.29σ signal it refused to call a discovery.
Game Autopilot

A concept to a store-ready build.

Seventeen stages cover design, art, code, machine playtesting, a verdict on whether it is actually fun, polish, and a store pack. It launches each stage itself.

Asked for the arcade game Gauntlet, it came back with HungerHall, twenty-four floors deep, and it plays.
Sneferu Coder

The coding superstars

An implementer, a pair coder, and a reviewer, working together on the same code. You pick the models.

The most capable coding system I've ever worked with. I'm not only the Sneferu.ai President, I'm also a client.
Adversarial Engine

Every system built by many models at once.

Sneferu’s adversarial convergence engine assigns independent models different responsibilities—author, adversary, reviewers, scorers, and judges. The makes disagreement part of the production process.

Every role is interchangeable with any model.
The Record

A record of every run.

Every run keeps its costs, decisions, failures, and recoveries. Work resumes after crashes, and a fleet of machines shares the load.

Lessons from old runs reach new ones without history being touched.
03How you drive it

Three ways to build.

Describe what you want and it writes the job. Fill out the form and launch it yourself. Or save it, put it on the board, and let it run while you sleep.

Run entry

One door into every production line.

Choose what you want out the other end, shape the run, choose the cast and spending limit, then read the exact request before anything runs.

Opening the page never starts or replaces a run. Nothing can launch from that state, and the button says so.
Snef

An assistant that genuinely knows what time it is.

The same assistant answering you in the console is the one reading every handoff between models before the work is allowed to move on. Ask it what happened to a run and it answers from the records, including when something is unavailable.

Snef understands the whole system, and sees every model handoff. Your go-to to build anything without using the UI.
Work Queue

Save a run, and the queue runs itself.

Any run can be saved as a task with its full configuration. Tasks can be reordered, queued, and left to start on their own overnight. Lessons from finished work are carried into new work.

Task queue. Easy way to turn ideas into products and systems without wasting any time.
04The workflows

Twenty-five production lines.

The workflows as of Aug 24 2026.

Set it and forget it.

The flagships
Business FoundryIdeate → pick → reframe → spec → code → test → fix. A fuzzy idea becomes a working product.Project roomIdea bloom
Game AutopilotA seventeen-stage autopilot covering concept, design, art, code, playtesting, and polish. The art is yours to redraw in the Asset Studio.Production
Research ExpeditionPut some models in a room together and make them do honest science.Findings
Goal MasterRuns one standing objective in rounds until its countable finish line is met or its budget stops it. Launches Research Expeditions and software runs until solved. Objective
Problem SolverFor questions with a right answer, it formalizes the claim, searches for a proof, and checks that proof with a program.Solved claim
The core
SpecProduces a written specification by argument: models draft it, attack it, and revise it until what survives is worth building. Check out the difference between the first and final draft. Paste the final spec into your favorite AI and ask, “This was produced from a single prompt. Any good?”ProductGame design
RefineThis phase takes a rough request and grounds it in reality, using reference codebases and attached files to build a clear, accurate problem statement for the spec writers. It deeply scans those sources to understand the existing system, constraints, and requirements before any money is spent building the wrong thing.
CodeBuilds software with the Sneferu Coder.Coder setup
UIBuilds a working interface, web, desktop, or mobile, on top of a system that already exists. Choose the framework in the launch window, or let Sneferu choose.Frameworks
DocsWrites the README, architecture, and API reference from the code as it actually is.Design doc
ResearchInvestigates a question and writes it up with real citations, keeping the places where its models disagreed.Read one
IdeateGenerates a wide field of ideas, scores them, and kills the weak ones. It is the brainstorm behind every product Sneferu builds.Read one
TaskMakes changes inside a codebase you already own: it writes the plan first, then does the work in your repository. This is what built Sneferu itself.
QuickOne draft, one review, done.
Quick FixA targeted repair: diagnose the problem, fix it, review the fix.
Code PipelineSpec → code → test → fix. You provide the problem.
Workflow DesignerTurns a plain-English description of a process into a new production line the system can actually run. Build a job that does whatever you want.
Spec FactoryTakes a finished plan and turns it into a queue of buildable pieces, then builds them one at a time, re-checking as it goes.
Asset ForgeGenerates a game’s artwork from its art direction, checks that every asset is legal to use, and puts it in the build.Asset studio
The helpers
Skill DesignTurns a wish into a new capability the system can reuse: packaged, tested, and discoverable afterward.
Improve (Scaffold)Takes a product that already exists and upgrades it in place, under the same checks as new work.
Playtest HarnessMachines play the finished game before any human does and report what they found.
Mac Build VerifierCompiles the game on a real machine, runs its tests, and launches it. This check turns “it should work” into “it ran.”
Science InstrumentsTwenty analysis tools that pull from the public archives of CERN, NASA, and USGS, plus telescope and genome databases.
Upgrade LoopLets the system improve itself: it proposes upgrades to its own machinery, then builds them as ordinary jobs.
05The operator experience

The operator experience.

Selected screens from the operator interface.

The adversarial tree: seven critics attacking one draft, competing rewrites, both scorers' marks, and every obligation raised against the work
The Adversarial Tree: the whole system, one screenThis is the argument itself, move by move. A draft sits in the middle. Seven critics approach it from different angles (vision, creativity, rigor, skepticism, novelty, feasibility, and synthesis), competing rewrites appear beneath them, and a judge chooses which one lives. On the right are both scorers’ marks with the spread between them, plus the obligations. Every problem raised against the work is marked addressed before anything is allowed to finish. Click any node and you read exactly what that model wrote. View a first draft beside the finished document
The Pipelines page: one job end to end: every model in every seat, both scorers' marks, the cost, and the full tree of drafts
Pipelines: the whole job, end to endThis page answers every question about a job at once. This one is a research run: Two-Detector Time-Slide Planet Hunt, hunting an exoplanet by borrowing the coincidence trick gravitational-wave observatories use. The cast: sixteen seats, each filled by a named model, so you can see who wrote, who attacked, who scored, and who settled it. The marks: both scorers on all six axes with their disagreements left showing: 85 against 88 on clarity, 86 against 82 on rigor. The money: $14.12 across 138 calls. And the tree: 84 nodes, every draft and every rebuttal, each one clickable to read the words the model actually wrote. The entire nine-turn, 15,240-word record is kept. It is the same run the assistant is reporting on in the next chapter.
The Run Room: a finished job retold from its own records: cost, duration, the critic's objection, and how far it got toward certification
The Run RoomRead any finished job like a report: what it cost, how long it took, what the critic objected to, and whether it earned certification.
The fleet page: coordinator plus three worker nodes with live runs, heartbeats, and code state
FleetThis screen shows the machines available to do work and what each is doing right now: one coordinator hands jobs to three workers, with a warning when a machine is running different code from the rest. Adding a node is as easy as attaching a Mac to your network and running a single install script.
Run Launcher: choose a flagship studio or workflow, shape the outcome, set the cast and controls, then review and confirm before launch
Run LauncherEvery job starts here. Choose a flagship studio, everyday workflow, or specialist system; shape the outcome; set the cast and controls; then review the exact request and explicitly confirm before anything launches.
Business project room: every stage down the left, a plain-English brief, and the running cost
Business Project RoomA business is built and watched here: every stage from market research to a shipped product appears down the left, with a plain-English brief saying exactly what is and isn’t waiting on a person, and the running cost.
Research expedition setup: expectations, data sources, what would prove it wrong, spending cap, and stop conditions, locked before spending
Research ExpeditionYou give it a topic. It splits the topic into testable directions, and you choose which ones to chase. For each one, it reads the existing literature and runs experiments against real public data. This screen is the setup: what it expects to find, which data it will use, what result would prove it wrong, the spending cap, and when to stop, all locked before a cent is spent.
Autopilot: master production lines on one page: the completed seventeen-phase game beside its siblings, spend against cap, truth panel
The autopilotGame productions are launched and watched here. Each one shows how far it got, what it has spent against its cap, which of its own checks failed, and includes a button to restart any stage.
The task queue: an ordered board the overnight conductor reads, with night-shift, finalizer, and ping toggles visible
TasksLine up the work, and the system runs it around the clock: whatever sits at the top gets built next, on demand or overnight while nobody is watching.
The Bloom Library: 6,639 judged ideas across 56 blooms, with scores, lineage, and launch history, with DrawSheet visible in the ranked list
Bloom LibraryEvery idea the system has ever generated and judged is here: 6,639 ideas from 56 brainstorms, each kept with its score, lineage, and objections. Any one of them can be sent straight into a build or a research run. The list shown is the brainstorm that produced the business product in the record.
Asset Studio library: materials and meshes with version history, qualification status, and license reasoning
Asset StudioA game’s artwork is produced and approved here: every texture and 3D model, all earlier attempts, the evidence that each asset is legal to use, and the votes that let it into the game.
Goal Master: a standing scientific objective with countable finish criteria
Goal MasterYou give it one standing objective and a finish line you can count. It breaks that objective into candidate missions. You choose which ones run, and it launches a full Research Expedition for each. When they return, it records what survived and what was refuted, counts that against the finish line, and plans the next round from what it learned. If a mission needs a tool the system doesn’t have, it writes a job to build that tool.
Problem Solver: a formalized claim searched under an oracle; failed branches marked dead, the surviving path certified
Problem SolverFor questions with a right answer, you give it a claim and a program that can check the claim. It writes the claim formally, breaks it into steps, and searches for a proof, running the checker on every attempt. Red branches failed the checker; the green path survived and earned a certificate. Maths stuff
Demo Studio: proof-directed film production; the screening room refuses to fabricate a preview; claims, rights, and privacy each carry a debt column
Demo StudioIt makes the demo video of a finished product, cut only from footage that was actually recorded during the work. It will not fake a preview, and every claim made in the film, plus every rights and privacy question, stays on the list until it is answered.
Usage cost evidence page: reconstruction labeled a floor, unusable evidence counted, per-provider daily bars
Cost AnalyticsThis screen shows what everything costs, in plain dollars: by job, stage, and model, over any window you choose. Every figure is rebuilt from the calls actually made.
Model Settings: each role assigned a provider and model with per-role fallbacks
Model SettingsThis is where you choose which provider’s model plays each part (the writer, attacker, or judge) and which model to use if one is unavailable. The roles are deliberately assigned across different providers, so nobody grades their own work.
Prompt Studio: all 418 prompts inventoried, each with version, hash, history, test bench, and usage
Prompt StudioEvery instruction the system gives its models is here. All 418 are listed, editable, and versioned, with a history of changes and a bench to test one before it goes live. None of it is buried in code.
How it works: the live map of the system: every subsystem clickable with a plain-language card, engineer notes, and its source files
How it worksA map of the whole machine you can click through. Every part opens a card that explains it in plain language, then in engineering terms, and names the exact source files that implement it.
The System Atlas map: documented nodes and relationships with drill-down doors
System AtlasThis is the deeper version of that map, built for technical diligence: every part of the system is documented against the code that implements it, with drill-downs into each one.
06The record

From its own intelligence to a playable game.

A few things Sneferu has built.

A
Exhibit A

It built the control system for its own model distillery.

One sentence became a system that finds where an AI model is weak, gives it focused new training, protects what it already knows, and refuses to replace it unless the new version is genuinely better.

asked for“Invent a novel fine-tuning method for optimal LLM performance in a target domain.”

Idea Bloom exploring 156 approaches to novel fine-tuning, with Failure-Directed Targeted Corpus Construction selected
The idea before the buildBloom explored 156 candidate approaches. Three survived the argument. The selected idea used a model’s failures to find the exact examples it needed next, instead of making a larger generic training set.
The receiptscode and bundled demonstration reviewed August 23, 2026
forged inMike’s Ideation run, then Sneferu’s Code Pipeline: spec, code, and test.
what it builtFDC Studio: software that studies where an AI model fails, finds the exact material it needs next, and decides whether the retrained version is actually better. The working system includes 22 Python modules and a seven-screen control room.
what connectsYou connect the model you want to improve, other models that judge it, search tools, human experts, and the training system you choose. FDC coordinates them, keeps the full record, and decides whether the new version replaces the current one. No single model or provider owns the loop.
how it worksStart with a fixed test. Group the mistakes that keep happening. Find or create lessons aimed at those exact weaknesses. Mix them with earlier examples so the model does not forget what it already knows. Train a candidate, test it again, and keep it only if it improves.
the safety rulesFDC will not let a model invent facts to teach itself. It will not use an example if it cannot trace where it came from. It will not replace the current model until the candidate passes a separate test it did not train on. If anything fails, the last good model and training set stay in place.
the demonstrationThe bundled demonstration runs three attempts against the same fixed 120-item test. One passes and is kept. Two fail and are rejected. FDC finishes with 36 approved training examples, each traceable to its source.
the recordIt keeps a history of every step. This demonstration records 33 events: 30 moves through the process, one accepted improvement, and two rejected attempts. Even the interface cannot go back and rewrite what happened.
testedWe ran 298 automated tests. Every one passed. They check how models are scored, how failures are grouped, where the training data came from, whether old skills are protected, whether bad candidates are rolled back, every screen, and the APIs beneath them.
The bundled demonstration uses simulated model responses so every decision can be tested without paying for a live training run. Real models, search tools, experts, and training systems plug into the same connections. What is proven here is the working decision system, permanent record, and ability to reject a bad replacement, not a completed production fine-tune.
FDC Studio showing three improvement attempts: one accepted, two rejected, and 36 traceable training examples
The decision screenFDC accepted the second attempt because it improved on a separate test. The third attempt made tool-use failures 24 percentage points worse, so FDC rejected it and kept the last good version.
FDC Studio permanent history showing each step and why the final attempt was rejected
The permanent historyEvery step is preserved. The ledger shows 33 events: 30 moves through the process, one accepted improvement, and two rejected attempts, including exactly where each failed and why.

Two more products from the same system.

Forged in: Business FoundryLife Skills Foundations started with a text from my mother about what school had stopped teaching her grandchildren. In 15h 53m, Sneferu returned a storefront, an owner’s console, seven printable packs, and five download bundles, all generated from one source.

Life Skills Foundations storefront: four printable modules, twenty-eight weekly lessons, twenty-five dollars one time
The storefrontFour modules, 28 weekly lessons, $25 once. The brief asked for stock-market lessons. The plan the system locked before writing any code excluded “investment recommendations of any kind,” and Module 4 now says concepts only on the page a parent reads.
Pilot Deck, the seller's console: build OK in 41 milliseconds, 257 tests passed, zero failed
Pilot Deck, the seller’s consoleIt also built the tool for running the business. It lets the owner edit a lesson, rebuild every pack, and track the pilot. We opened it on a machine that had never seen it and told it to rebuild itself: build OK in 41ms, 257 of its own tests passed, none failed.

Forged in: Business FoundryRestoration Copilot is software for a classic-car restoration shop: take the car in, identify every part it needs, find those parts, and show the shop how to install each one. It serves a different trade, but comes from the same production line and runs on Sneferu itself.

Restoration Copilot: a 1969 Camaro job with intake, parts manifest, budget, sourcing, installation guidance, and a 3D-model stage
A job on the floorA 1969 Camaro is open now, with intake, a parts manifest, a budget, sourcing, installation guidance, and a 3D-model stage. Finding a part is not the finish line: the job also shows the mechanic how to install it. The machine’s own spending is capped at $100, and it raises a flag the moment it needs a human: an ambiguous part, a budget overrun, or a dead end in sourcing.
The intake bay: photograph the car, log the parts you already know are bad, then seal the intake to start the scan
The intake bayPhotograph the car, log the parts you already know are bad, then seal the intake, which is what starts the identification scan. Sealing is deliberate: the scan runs against exactly what was sealed, so what it was given is never in question afterward.
B
Exhibit B

From a single prompt to a bug free playable Steam game.

Seventeen stages ran without a person in the loop, and the machine that decided whether it was any good had to play it first. It was surprisingly good and bug free.

Human art in a game development loop is my preference, but to prove the system everything must run safely and without a human first.

asked for“Something like the arcade game Gauntlet, from the ’80s.”

The receiptsAugust 15–16, 2026
forged inGame Autopilot, Sneferu’s 17-stage game production line.
what it didIt handled concept, design, code, machine playtesting, a review of whether it was actually fun, polish, and the store launch pack. All seventeen stages finished.
buildPLAYABLE: A machine played the finished game, and the recording of that session is kept.
post-runA 15-entry defect audit was published against our own run: 13 defects were confirmed against the code, and one was reclassified with the correction printed.
Not claimed as a shipped game: audio awaits the founder’s own soundtrack, and two of its own checks are still unproven. The complete production ran from start to finish, with its defect list published.

Not gonna lie, this was the hardest system I’ve ever built. I’m still shocked it works. I wouldn’t build it again. It was a fight, but Sneferu came out of it with an immune system. Hands-free game development won’t always be smooth—and probably shouldn’t be.
The Autopilot console showing the game master complete, spend against cap, and every failed gate named
Working screenThe game’s console shows the completed production and a panel listing every check that failed instead of hiding them.
Run history for the seventeen-phase game: per-phase status, gate decision, and cost
Working screenThis screen shows the same production, stage by stage: what each stage cost, what was decided, what is still unproven, and a button to restart any of them.
0:00 / 3:27
The build, playingHungerHall is what came back from the one request at the top of this exhibit. The recording starts at the title screen and moves into the dungeon: seven floors down, enemies, doors, and a health bar that is also the clock. The document opened at the end is the design document the system wrote for itself at stage two: the engine, the price, the audience, and the rule that it may not lift a single name or sprite from the arcade game it was asked to answer. View the build specification, first draft and final
C
Exhibit C

It found a signal at the discovery threshold and refused it.

This instrument retrieves raw public data, runs the statistics itself, and knows the difference between a finding and a fluke. It was pointed at nine archives where the answers are already known to test whether it could be trusted on archives where they are not.

asked for“Run a cross-archive calibration gauntlet: nine independent public archives.”

The receiptsreal public data
forged inResearch Expedition, Sneferu’s production line for testable questions, data wrangling, experiments, and findings.
particle physicsFrom raw CERN Open Data: J/ψ at 265σ, the Z boson at 201σ, and five more known particles, each at the confidence level physicists call a discovery.
exoplanetsFive known planets were recovered from raw telescope archives, including a rocky-planet signal of 147 parts per million, near the limit of the instrument that found it.
earth & spaceThe earthquake-frequency law was recovered from 15,417 real quakes, along with the 11-year solar cycle, the lunar tide, and all four known gaps in the asteroid belt.
the refusalsREFUSED: A 5.29σ bump its own statistics couldn’t defend, plus a seismic anomaly rejected because the literature itself urged caution.
VerifyResults and experiment software can be sent to experts in their field with confidence the research is not a hallucination of some kind.
Nothing above is claimed as new. That is the point: an instrument is only worth aiming at an open question once it has been made to prove itself on a settled one. It passed, and then it refused a result at the threshold physicists call a discovery, because its own trials correction would not carry it.
See one for yourself  →  sneferu.ai/research/rex-c8661466
A live research run in the tree engine: seed drafts from independent models, critiques, scorer pairs and judge decisions
Working screenThis is a live research run: seven independent first drafts from five different makers, structured critiques, two scorers and a judge per turn. Disagreement is on screen because disagreement is the data.
The idea bloom: 200 generated research directions clustered by strategy, with survived, mutated, and killed counts
Working screenThis screen shows a 200-idea research bloom with its count printed: 39 unchallenged, 10 survived, 63 mutated, 63 evolved, 25 killed. Each petal opens to a full testable brief.
The scored idea list with rubric weights and per-idea provenance
Working screenThis screen shows the same ideas scored, with the rubric and its weights printed beside the list, each survivor showing what it grew out of and how it changed.
08Why now

The models got capable. The finish line didn’t move.

Models keep getting better. The hard part is still getting from a prompt to finished work.

09Founder & contact

Mike Lange

A word from the founderMike Lange on why he built Sneferu and what he wants it to make possible. 4 minutes, 26 seconds

Hi. I'm Mike. I've spent nearly thirty years building technology.

I've built everything from kitchen management software for the CIA - Culinary Institute of America, bedside charting and diagnostic tools for er doctors, to custom tools, strategies and indicators for NinjaTrader. And a thousand things in between.

Today I live in a farmhouse in Havana, Florida, with three cats, and nature all around me. See, I like learning my way into things and making something real. I never just dip my toe in the water. I've known about LLMs since the paper describing the concept was written.

So when they said, five years later, "Here you go, here is this plethora of maths that will do everything you can do AND take your job", I was nervous and excited. I'm so sick of working. But then, when it couldn't do my job and still can't, I was quite offended.

So I built a system that CAN do my job, and theirs too. I say that with love.

Sneferu is my version of what AI should be.

For better or worse, this is the world we've been given and the tech to run it, let's make the most of it, try to keep it honest, reinvent if need be, and do some actual good.

Contact me

I am looking for a co-founder, an early investor and folks in the academic/research world who would like to collaborate and solve some real problems. If that is you, I’d love to hear from you.

Sneferu is not available to the public. The forms it will be available have yet to be determined.



One reply, from me. No mailing list.

10Roadmap

On the horizon.

01
Sovereign intelligenceModels for both Snef (the brain) and Sneferu Coder, built from run data.
02
Vertical CastDeep domain intelligence for a specific industry or purpose, trained on the failures that matter.
03
Low-voltage personal manufacturingPCBs, materials, 3D printing, and self-sufficiency.