Verified software optimisation · October 2026
Software that proves its own improvements.
From bare metal to a proprietary, machine-optimised stack. Every gain verified before it counts.
- Cost per request
- −30.9%
- Real bugs fixed, same model
- +75%
- Planted cheats caught
- 31/31
Why now
Compute is one of the world's largest bills, and much of it is waste.
$723B[1]
Worldwide public cloud spending, forecast for 2025.
29%[2]
Share of IaaS and PaaS spend that organisations estimate is wasted.
~$725B[3]
Capital spending planned for 2026 by Alphabet, Amazon, Meta and Microsoft combined.
How the giants do it
The largest companies already prove that automated optimisation pays.
Google DeepMind · 2025
0.7%[4]
of Google's worldwide compute recovered, on average, by an AI-discovered scheduling heuristic.
Meta · CGO 2019
Up to 8%[5]
faster data-centre applications from a binary optimiser, on top of existing compiler optimisations.
Each win is one technique on one layer, built in-house by a company that can afford a research team. Colloid works across every layer, for everyone.
What Colloid does
It improves real software, and keeps a change only when it is proven.
01
Improve
Colloid finds changes across every layer of a running system: code, data access, runtimes, compilers and operating-system settings.
02
Prove
An independent verifier checks every change and measures its gain with statistical confidence. Where it is decidable, the change is proved formally.
03
Keep
Every verified gain is kept with its evidence: the proprietary raw material for rules, and for the stack we are building.
The vision
From bare metal to our own stack.
Every layer of a computer system is a place to find verified gains. Today the engine works from the operating system up to application code. The destination is a stack of our own, machine-optimised and proven at every layer.
Verified results
Services & applications
Service code, and fixes to real open-source issues.
- 75% more real GitHub bugs fixed than our first engine generation, with the same AI model.
- The first fixes on issues created after the model was trained, so no answer could be memorised.
- Graded by each benchmark's official harness, not by us.
Select a layer to see where it stands. Status as of October 2026.
Verified results
Every number is pre-registered, measured and tested.
−30.9%
cost per request on a production-style web service
95% confidence interval: 27.3% to 34.5% lower. On a hidden holdout workload the search never saw, the gain was 26.6%, and median latency fell 46.9%. Every answer stayed identical, checked independently.
- Baseline = 100
- Optimised by Colloid
Real-world bugs
Same AI model. 75% more real bugs fixed.
On 30 issues from SWE-bench Verified[6], the industry's standard test of fixing real GitHub bugs, our fourth engine generation resolved 21 against 12. The model, the verifier and the time budget were held fixed.
Exact McNemar p = 0.012, pre-registered before the run.
No memorisation possible
On bugs fixed after the model was trained: from 0 to 7.
Old benchmark issues can leak into a model's training data. So we tested on issues created after the model's release. One change to how the engine searches took it from 0 to 7 of 40. Exact p = 0.016, none lost.
Reported in full: On older issues the same change did not move the score (20 against 21 of 30), within the run-to-run noise we measure: two identical runs differ on 2 of 30. We publish null results too.
Rigour is the product
Pre-registered, tested and published, failures included.
10
experiments pre-registered before they ran
31/31
planted cheats caught by the verifier
7/7
database optimisations formally proved before running
2/30
our own measured run-to-run noise
Not yet proven, and we say so: That knowledge from one system speeds up the next. Its first test did not pass, and a stronger re-test is designed.
The moat
A growing record of what provably makes software better.
Rarer than code
Each record holds a change, where it applies, and the evidence that it worked, or that it did not.
Private by design
Customer code stays under the customer's control. The evidence and the rules distilled from it stay ours.
Built to compound
Records seed future searches, become rules, and in the end become the stack itself.
Competitors can copy a technique. They cannot copy years of verified evidence.
Who it helps
Anyone whose software runs at scale.
Hyperscalers & big tech
Fleet-wide efficiency at Google scale, where a fraction of a percent moves budgets, and every change ships with proof that it is safe.
Cloud-native companies
Attack the 29% of cloud spend that surveys call waste, without risky rewrites.
AI infrastructure
Verified efficiency for the services around AI workloads. Accelerators are on the roadmap.
Software vendors
Fixes to real bugs, graded by the project's own test suites before anyone reviews them.
Regulated industries
Banking, health and government get an auditable record of correctness and gain for every change.
Platform & SRE teams
Continuous tuning of databases, runtimes and operating-system settings, with the evidence attached.
Milestones
Each step is gated by evidence, not by a calendar.
Achieved · 2026
- −30.9% cost per request, verified
- One engine for Python, C, Go and TypeScript systems
- 75% more real bugs fixed with the same AI model
- First fixes with no memorisation possible
- Formal proofs for database optimisations
Next · Months
- Self-correcting search on long tasks (in progress)
- Optimising real third-party repositories
- Rules that work on systems they were not learned on
- Formal proofs beyond the database
Destination · Year+
- Our own language and compiler
- Verification across x86, ARM and accelerators
- A complete proprietary stack, every component traced to evidence
Colloid
Proof, not promises.
We are looking for investors and design partners that run software at scale and want every efficiency gain delivered with proof.
Get in touch