AI & machine learning research

Research that survives its own controls.

Hatteria Labs studies how language models work from the inside. We publish what we find — including the results that failed — and build products on the parts that hold up.

We measure

Empirical work on model internals: memory compression, weight structure, quantisation, hybrid architectures. Every claim carries a null control it has to beat.

We publish

Methods, thresholds and negative results, in the open. A rejected hypothesis is a finding — it is where most of the value in this work lies.

We build and operate

Three products in production, built and run by the same people who do the research. What we learn goes into software people pay for.

Products

Software we build and run

Each product is live, paid for by its own users, and runs on our own servers in the European Union. They fund the research and keep it anchored in what works outside a benchmark.

More about the products →

torumata.com

An audit for the era of AI search — see how language models actually read your site.

What it does

  • A readability score for AI, broken down by category, so progress is measurable rather than felt
  • An analysis of which questions the site can actually answer, run through a language model
  • A site tree with flags marking the pages where the problem sits

For: Site owners and marketing teams whose traffic is moving from search results to chatbot answers.

AIDJ

Live

aidj.cloud

The DJ that runs your party on autopilot.

What it does

  • Guests request a track from their own phone after scanning a QR code, with nothing to install
  • A spoken DJ introduces the track in its own synthesised voice and crossfades into the next one
  • When requests dry up, the queue keeps filling itself from the mood of the event, so the room never falls silent

For: Weddings, company parties, bars and clubs, birthdays and school events — anywhere the music matters and nobody should have to spend the evening managing a playlist.

slidify.cloud

Your guests' photos, live on the screen.

What it does

  • Guests scan a QR code and upload from a mobile browser — no app, no account, nothing to explain
  • Photos reach the projector or television within seconds of being taken
  • The slideshow skips what it has just shown and favours pictures nobody has seen yet

For: Anyone hosting a wedding, a celebration or a company event who would rather collect the evening's photographs as they are taken than retrieve them afterwards.

How we work

Rigour is the product

Most of what looks like a result in machine learning is a measurement artefact. Our process is built around catching our own.

Every number has a null beside it
An observation is printed next to the control that randomises the structure it claims to depend on. If the null keeps up, the claim fails — and is recorded regardless.
Thresholds are set before the measurement
Decision criteria go into the docstring ahead of the run, as named booleans rather than prose. A verdict written after seeing the data is not a verdict.
Corrections stay visible
When a later replication overturns an earlier conclusion, the correction is published next to the original. Three of our own conclusions have been withdrawn this way, one of them an hour before release.
A proxy measures the proxy
A quantity indexed by availability measures availability, not value — until value is measured directly. We hit this four times in four different domains, each time with a finished story that would have passed review.
The end-to-end effect is the only score
Captured variance, reconstruction error and attribution scores all mislead. We measure what the change does to the model itself, on held-out data.

Working on something adjacent?

Alongside our own products we accept a limited number of research and engineering engagements — typically where a question needs to be measured properly before anyone builds on the answer.

Get in touch