how it works

Not a data feed. A read on what happens next.

Anyone can buy a copy of the register. Scout is something else: all 5,695,465 UK companies indexed with their filed history held point-in-time, the 1.5M-company scored universe watched live, enriched from every tier of source we can reach — open registers, credentialed feeds, and proprietary layers we built ourselves — and run through models that put a calibrated probability on the things you actually care about: whether a business is heading for a sale, for new borrowing, or for a refinancing.

Real time, whole market Every UK company, not a curated slice — and live: a scored business files, changes hands or takes on debt, and its read is recomputed within minutes of the register publishing, not on next month's data drop.
Machine learning Models trained on years of real outcomes and calibrated so a stated 10% means close to 10% of that cohort really did it — not a score out of a hundred.
Named patterns 43 hand-built archetypes — succession windows, cash fortresses, quiet compounders — the buyout shapes measured against the backtest, so you can hunt a shape and not just a score.

The index

Where does the data come from?

We don't resell anybody's database. Scout is an index we assembled ourselves, and it draws on every tier of source available on the UK market:

  • Open public records — the statutory registers every UK company is legally obliged to file into, plus insolvency notices, public contract awards and property records. This is what makes coverage a census rather than a sample: the filing obligation means no company can be invisible to us.
  • Credentialed feeds — free to access but not bulk-downloadable, which is why almost nobody indexes them: real-time channels and per-entity interfaces that carry depth the bulk files never contain.
  • Our own proprietary layers — and this is the part that cannot be bought from anyone. A machine read of the narrative text inside filed accounts, not just the tagged numbers; an identity graph that resolves the same operator across the companies they touch, wherever the register's evidence allows it; and seven years of point-in-time history we rebuilt and keep ourselves so the past can be replayed honestly.

How those tiers are ingested, reconciled and kept in sync is our own engineering, and we keep the recipe in-house — it is a meaningful part of what makes the index hard to copy. What we will always be explicit about is the provenance of any individual figure we show you: every number in a dossier can be traced to the filing or the observation it came from — and where a figure is our estimate, it is labelled as one.

What data do you hold on each business?

Financials come from the tagged accounts every company must file — so we have real balance-sheet series, not estimates:

30.3M
annual account rows parsed
6.2M
companies with filed history
2016
history reaches back to
  • Financial — net assets, cash, creditors, trade debtors, stock, fixed assets split into land & buildings / plant & machinery / vehicles, turnover, gross profit, profit, and average headcount. Plus derived measures: gearing, current ratio, Altman Z, Zmijewski and Taffler scores. Every figure is stored with its register date, so the past can be replayed.
  • Non-financial — ownership from the PSC register including the owner's birth month and year (so, their age — the single most under-used field in UK company data), corporate ownership chains up to the ultimate parent, directors and their appointment history across companies, charges and who holds the security (a high-street bank versus a private lender is itself the credit tell), filing punctuality, sector, incorporation date and region.
  • The text nobody indexes — for every company you unlock, an LLM reads the actual notes and audit report inside the filed accounts and extracts what the tagged data cannot carry: material going-concern uncertainty, the audit opinion, director loans and their direction, post-balance-sheet events, customer concentration, restatements.
  • Beyond the statutory filings — insolvency notices, winding-up petitions and property holdings, joined onto the same company record so one profile carries all of it.

Headcount deserves a note: small companies are exempt from disclosing turnover, but nearly all must disclose average employees. That makes the staff trend one of the best free proxies for revenue across the long tail — and one the models genuinely use.

Why don't you print an estimated-revenue figure, when everyone else does?

Because a single number would be a guess dressed as a fact. Small UK companies are exempt from publishing a profit & loss account, so turnover is simply absent for the overwhelming majority of the register — and the ones that do publish it are the big end of the market, not a sample of the rest:

~97%
of live companies never disclose turnover
×1.4
median error of our best estimator
×5
the widest band we will ever print

We built the estimator anyway — a gradient-boosted model over the full filed history, split by company so no firm's own past could leak into its own prediction — and measured it. Even at its best, the median estimate is off by a factor of 1.4, and nearly ×1.9 on micro-companies. A company labelled "£2.4m" that actually turns over £600k sits comfortably inside that error — and one such phone call would rightly cost us your trust in everything else on the page, including the probabilities, which are validated. The limit isn't the algorithm: most of the signal lives in the P&L history that the abridged regime omits. You cannot model information that was never filed.

The industry prints a confident single figure anyway. Every major database sells "estimated revenue" set in the same typography as a filed fact, with no error bar and no tier — and the error we measured rides along silently, call after call. Our best estimator, holding the company's full filed history, is off ×1.4 at the median; a sector-ratio without that history does materially worse. Confidence is not accuracy.

So the dossier does the honest version instead. Where turnover is filed, you see the real figure. Where the filings support an honest range, we print a band, labelled as an estimate, with the tier it came from. And when the honest band would be wider than ×5 — as it is for most micro-companies — we print nothing at all, because a range that wide is not information. What we never print is a confident single number.

The precise figures a company genuinely files are always there: net assets, cash, trade debtors, stock, fixed assets and headcount, as a dated series with the year-on-year moves — and "files abridged accounts" is itself a fact about size worth knowing.

How current is it, and what does "real time" actually mean?

A persistent connection to the register's filing stream, running around the clock. The stream delivers the event within seconds of publication; we then re-fetch what changed, recompute the company's patterns and re-score it — typically minutes between a filing being published and the index reflecting it, not next month's data drop. The monitor shows the stream running, live.

Underneath, each layer of the index refreshes on its own schedule — some daily, some weekly, some monthly — so financial depth and ownership stay current without waiting on a quarterly rebuild. Live re-rates are journalled from the day they happen, and today's model is replayed over the point-in-time vintages, so you can see how it would have read the company each year. Momentum is often the more interesting read.

The models

What exactly do the three scores predict?

Each score is the probability of one specific, checkable register event, inside a stated window. Not a mood, not an index out of 100 — an event you could verify yourself a year later:

  • Sale likelihood (~24 months) — an external acquisition: control moving to a corporate owner where an individual held it before. It happens to about 0.33% of scored companies in any 24-month window.
  • Credit appetite (~12 months) — the company takes on new secured borrowing: a fresh charge appears on the register. Base rate about 1.1% a year. It is an appetite for capital — growth, an acquisition, a refinancing — and deliberately not a distress or solvency judgement.
  • Refinance (~12 months) — an existing charge is repaid and replaced inside the window: the classic refinancing signature. Base rate about 0.24%.

One honesty note we'd rather volunteer than have you discover: credit and refinance overlap heavily at the very top — roughly two thirds of their top-1% is the same companies, which is natural (a refinancing is new secured borrowing). The sale model is genuinely independent of both.

Because every label is a register fact, training data is replayed from filings — not surveys, not panels, not self-reporting.

How honest are the probabilities?

The models learn from years of real outcomes and are tested strictly walk-forward: judged only on periods they were never trained on, so nothing leaks backwards from the future. It's the discipline most vendors skip, because it is the one that exposes a weak model.

Probabilities are then calibrated on a full held-out year: the register is cut into buckets by score, the real outcome frequency of each bucket becomes the scale, and the curve is never extended past its own data. The one adjustment that can lift a probability above the curve's ceiling was itself measured on a holdout — promised 33%, realised 32%. When Scout says 10%, close to 10% of that cohort really did it.

0.33%
base rate for a sale in 24 months
1.1%
base rate for new secured borrowing in 12 months
2018–25
point-in-time vintages behind the tests

That first number is the whole point of the product: sales are rare. A card reading "≈13%" is not a modest number — it is roughly 40 times the event's base rate, well over an order of magnitude above the register. Every score also ships its drivers, so you can see which facts moved it and decide whether you agree.

And it stays a signal, not a verdict. These are cohort estimates about a population, never a claim about one business's intentions.

Have the models been tested on data they never saw?

Yes — and it is the test we would show a sceptic first. We rebuilt the register exactly as it stood on 30 June 2025, a vintage the models were never trained on, scored every company blind, and waited for the 12-month windows to close (the sale window is still running — see below). Then we opened the envelope:

ModelBaseTop 0.1% — actualPredictedLift
New secured borrowing, 12m1.08%72.4%69.1%67×
Refinancing, 12m0.24%30.6%30.4%127×
Business sale, 24m0.33%12.0%17.2%36.7×

Two things matter more than the loud multiples. First, the predicted column: the model said 69.1% for its top credit slice and 72.4% happened; said 30.4% on refinancing and 30.6% happened. It doesn't just rank companies — it prices the probability, and the price is honest. Second, the bottom half of all scores produced near-zero events: knowing where not to spend your week is half the value.

The sale row is a floor, not a miss — its label runs 24 months and only 13 had elapsed when we measured. Nothing was retrained or cherry-picked after the fact: one frozen model, one unseen year, one readout. The same row is printed on the monitor and changes only when the validation is re-run; the full method is in the write-up.

How do I read a card?

Four elements, each meaning exactly one thing:

  • ≈N% — the calibrated probability of that event, in the stated window. The ≈ is honest: it is a cohort estimate, and it is capped — we never print a probability the held-out data couldn't support.
  • — the same number as a multiple of the scored-index average, because base rates this small defeat intuition: ≈13% on sale reads modest and is roughly 40× the field. "≈ avg" means nothing notable either way.
  • top N% — where the company ranks, for that signal, among the ~1.1M companies that carry a model read. A position in our ranking, not a fact about the business.
  • ▲/▼ pp — momentum: how the sale read moved on the latest re-score, in percentage points. Moves too small to mean anything are suppressed.

And when a card says "no model read yet", that is the honest state: too little on file to score. We would rather show you the gap than a made-up number — the same reason we don't print estimated revenue.

What are the patterns, and how are they different from the score?

A score ranks. A pattern explains — and it is something you can point at and hunt. Each of the 43 archetypes is an explicit set of conditions over the register: Succession window (retirement-age owner, solvent, real trading company), Cash fortress (cash over half of net assets while headcount stalls — the owner is converting the business into liquidity), Quiet compounder, Never borrowed, Hidden distress, and so on.

Every buyout archetype carries a measured lift from the walk-forward backtest, so you can see which shapes actually precede a sale rather than merely sounding clever (the newer credit-side patterns queue for the same harness). They are recomputed the moment a company files, and entering a pattern is a dated, journalled transition — the origination moment.

The same idea runs over people: a dozen owner archetypes (serial exiter, accumulating, winding down…) so you can spot the operator who is quietly selling down a portfolio before any single company of theirs looks interesting.

Why is there no failure score?

Because we built one, measured it, and it was a lie with excellent metrics. The obvious label — "company disappears from the register" — turns out to be roughly 95% routine housekeeping: dormant shells voluntarily struck off. A model trained on it scores beautifully (AUC 0.93) at predicting… which empty shells get tidied away. Genuine distress is a few percent of that label, drowned in noise. We dropped it rather than sell it.

The deeper reason: insolvency is a fork, not an endpoint. Administration and a CVA are rescue procedures; many distressed companies refinance, restructure or sell rather than die; a members' voluntary liquidation is a solvent owner's exit — closer to a sale than a death. A single "failure risk" number flattens all of that into a scare.

So Scout does two honest things instead. Procedures and winding-up petitions are shown as dated register facts — the petition being the earliest public marker, weeks before the register shows an order. And the credit score stays what it says on the tin: an appetite for capital, never a solvency judgement.

Names, fairness & privacy

Why do the score cards hide the company name?

Because the join is the product. Names, filings and register facts are public and free — here and everywhere. The models' read is ours. What an unlock buys is the two joined together, so the anonymous board shows the read in full and holds back the name.

The same discipline runs the other way: a signed-out page never pairs a name with our read, and named lists never sort or filter by score. It is the business model and the defamation firewall in one rule — an exact score order over named companies would publish the scores as a permutation. It is also why a signed-out card carries neither name nor number: a company number is one free register lookup away from the name, so withholding one without the other would be decoration.

I set a narrow filter and got nothing — why?

The anonymised view refuses to be that specific, on purpose. Two guards are at work:

  • The cohort gate. A filter combination matching fewer than a few dozen companies returns no cards — and the counter reads zero rather than "3 match", because an exact count over arbitrary filter intersections is itself a way to fingerprint companies without ever seeing a card.
  • The look-alike rule. A card is shown only when enough similar companies exist that its scores can't be pinned on one of them — either the cohort is large, or the scores inside it genuinely disagree. Roughly 5% of cards fail that test and are quietly skipped; we measured the cost of skipping them at 0.0% of headline quality.

Sign in and the named board answers precisely — names are free with an account; it is the name-plus-read join that stays paid.

Is this legal, and what about people's data?

Everything here is computed from official public records — principally Companies House, which is published under the Open Government Licence, alongside other public sources. Owner names and ages come from the PSC register, which exists precisely so that company control is public.

That said, they are still personal data, and we treat them that way: every public page that carries model output is anonymised and banded with a minimum cohort size, so a filter can never single out an individual business; register facts — filings, procedures, appointments — are shown as-is, because they are already public. We keep the provenance of every figure. Nothing here is an automated adverse decision about anybody, and the framing stays "worth a closer look", never a verdict. The detail lives in the privacy note and the terms.

Using Scout

What's free, and what does an unlock actually buy?

Free, forever: the whole board with precise register facts (names appear on sign-in), search, the pattern library, the monitor's dashboards, the journal, the blog — and the weekly pack below. Signing up also grants a bundle of unlocks to spend.

An unlock is permanent and opens everything we hold on one company: the name beside the three model reads and the facts that drive them, the machine read of the accounts' narrative text (going concern, audit opinion, director loans), a written lead brief arguing the company honestly as a buyout and as a credit target, and live monitoring — every register move lands in your feed from then on. Unlocking the same company twice costs nothing.

One note: the very top of the index — top-1% on any signal — is desk business. A near-certain event is deal flow, not a lead, and the strongest cards may open through a conversation rather than a token.

How does the weekly free pack work?

/start asks two questions — are you buying or lending, and which sectors — and deals you the week's ten per sector: two anonymised cards from the very top of that ranking, and eight named companies stepping down a ladder of bands, so the numbers fade card by card exactly as the ranking does. Signed in, every dossier in the week's pack is open in full until Monday deals the next one; unlocking is only needed to keep one for good.

The draw is fixed for the week — a reload can't re-roll it, and everyone picking the same sector sees the same ten. The named giveaway never contains a top-1% card; the pair at the very top stays anonymised. It is a shop window with the discipline of one.

What happens after I unlock — what does monitoring mean?

Every company you've unlocked reports to your feed: filings the day they land, procedures the moment they're on record, and re-ratings — but only when the read actually moved; a re-score that changed nothing isn't news.

Unlock a person and their whole portfolio starts reporting register facts — named, because facts are public. What stays sealed is our analytics on the portfolio companies you haven't unlocked: you'll see how much moved across the portfolio, week by week — never which company carried the signal. The count is free; which is the paid read.

What is the insolvency register on the board?

A different population, one chip away: every UK company with a live insolvency procedure — administrations, liquidations, CVAs, receiverships, with the practitioners on record — plus winding-up petitions from The Gazette. The petition matters because it is the earliest public marker: Companies House only ever shows the end of that road, the order, weeks later — if it happens at all.

Model scores deliberately don't ride there: a company already in a procedure needs a practitioner list, not a probability.

Can I get a list built to my own thesis?

Yes — that is the main way people use it. Tell us the shape you hunt (sector, size, region, owner age, balance-sheet posture, whatever defines your mandate) and we encode it as a pattern, run it across the whole index and its history, show you the backtest, and keep it re-screened live exactly like the standard set. New matches reach you the day they qualify.

Use the brief form, or reach Elijah Podavalkin on LinkedIn.

What does it cost?

Signing up is free and comes with a bundle of unlocks — enough to open real dossiers and judge the read for yourself. Beyond that, today unlocks come as keys: bought directly from us, or redeemed from a promo code on your profile. Self-serve top-ups are on the way; until then, tell us what volume you need and we'll sort it the same day.

Who is Scout built for?

Buyers — search funds, micro-PE, first-time acquirers. The best targets are grey mice: profitable, established, an owner past sixty, and nothing happening — which is exactly why event-driven tools never surface them. A census can't miss them.

Private credit — funds and brokers reading appetite and refinancing timing: who is about to borrow, whose facility is nearing its end, which recapitalisations are forming.

Advisors — accountants and corporate-finance people who want to be the first call. Pattern entry is an origination moment: the week a company starts looking like a succession sale is the week to phone the owner.

Want a specific number of companies, to your own brief? Tell us the shape and the volume — Elijah Podavalkin comes back to you personally, usually the same day.
Request a brief