What it is
What does the tool actually do?
It reads the documents a business already produces — board packs, budgets, org charts, hiring plans, pipeline — against the strategy its board approved, and surfaces where the two disagree. Every finding carries a verbatim quote from your own documents on both sides: the claim, and the observation that appears to contradict it. Nothing new has to be written to use it.
Is this AI writing our strategy?
No, and it never will be. The strategy comes from your leadership, captured in a short structured interview. The tool holds the organisation to it. The machine only ever proposes candidates; a named person decides which become conclusions, and the innocent explanation for every candidate is considered on the page before anyone acts on it. In future the tool will also help you state your strategy as real choices — drafted as competing options, with the evidence and resources each would need — but it will not make them for you. Ever.
How is this different from pasting our board pack into a chatbot?
A general assistant is fast, and it is also unaccountable. It has no structured model of your strategy, no requirement that every quote appear verbatim in the source, no second model checking each inference blind, no memory of what was decided last quarter — and your documents leave the building. A hallucinated contradiction in a boardroom is not a time-saving. This tool is built so that a finding is either evidenced, checked and owned, or it is not shown.
What does a finding look like?
One line, two citations, and an exit. For example: "Priority says retention before acquisition; 78% of Q3 marketing spend is acquisition" — with the claim cited to the leadership interview, the observation cited to a page of the budget, an innocent explanation (a one-off launch campaign, perhaps), and what would clear it. Findings that survive scrutiny become part of the record; findings that do not are cleared, visibly.
Trust and accuracy
How wrong can it be?
Less wrong than you would reasonably fear, and honest when it might be. Three mechanisms: every quote must appear verbatim in the source document or the finding is not shown; every inference passes a blind second opinion — a separate model, shown only the two cited passages, must agree the reasoning holds, and that check can be pinned to a different provider entirely, so it does not inherit the first model’s blind spots; and everything is presented as a candidate with its innocent explanation attached, so a wrong candidate costs minutes, not credibility. Disputes are recorded, and a disputed finding is visibly downgraded.
What if the finding has a perfectly good explanation?
Then it becomes an exception — recorded, with an owner and an expiry date. That is a feature, not a failure: an explained deviation this quarter that is still there next quarter, after its explanation expired, is exactly the kind of quiet drift the tool exists to catch. Explanations are believed once, and re-tested afterwards.
How do we know a report has not been quietly altered?
Because it can prove itself. Every report closes with a document register — the exact filenames, versions and integrity checks of the documents it was built on — and every issued PDF carries a signed fingerprint recorded at the moment of generation. Upload any copy back into the tool and it answers plainly: authentic (with when and for whom it was issued), altered, or unknown. The same discipline applies on the way in: source documents keep their exact bytes and hash checks, so what a report cites can always be re-produced and verified. If sections are removed for wider circulation, the report says so — what was omitted, why, by whom — as part of the document itself.
What if our strategy cannot be tested at all?
Then that is the first finding, and it is the most useful one. Statements no evidence could ever contradict — "quality is in our DNA" — are flagged as untestable rather than quietly ignored. Many organisations discover this is true of most of their strategy. The fix is short, structured work (a testability review), not a transformation programme, and a team that stops there has still learned which of its strategic statements can be shown to be followed. Try the readiness check to see where you stand.
Confidentiality and security
Where do our documents go?
By default, nowhere. The tool runs as a dedicated, non-shared instance: on a laptop, in your own Azure or AWS environment, or hosted by us in a boundary dedicated to you — one instance per customer, never shared infrastructure. In cloud deployments, models are served inside your own cloud boundary over private endpoints, so documents never leave it. In every deployment, sensitive names and figures are pseudonymised before any model sees them.
Can it be run in the cloud?
Yes, and you can choose who runs it. Either it sits in your own AWS account or Azure subscription — your bill, your encryption keys, your identity provider, with our deployment access granted, logged and revoked by you — or we host it for you, in an account dedicated to your organisation alone. Not a tenant in a shared system: one customer, one boundary, one set of keys, whichever route you take. The way out is the same either way — a portable export of everything, available any day, that restores on the other cloud, on our hosting, or on a laptop. If hosting it yourself is not what you want, talk to us and we will run it for you. The four options, side by side.
Which model does it use, and can we choose?
You choose, per instance. In your own cloud it uses the model service inside your boundary — Claude on Amazon Bedrock, or Claude on Azure AI Foundry — reached over a private endpoint. On a laptop it can use the Anthropic API with your key, or the Claude desktop subscription you already pay for. It also runs against OpenAI, or against an OpenAI-compatible gateway you already operate inside your own network. Options your environment cannot reach are shown greyed out with what is missing, so nothing is hidden behind a support call.
Can it run with no external model at all?
Yes. It runs against a private model on hardware you own, through Ollama, and then no document — pseudonymised or otherwise — makes any network call. An engagement can be marked local-only, in which case the tool refuses to use a hosted model rather than quietly falling back to one; if the private model is not running, the work stops instead. The trade is capability: a model on your own hardware is smaller than the frontier ones, so it is the right choice for the most sensitive material rather than the default for everything.
Where do the API keys live?
Where your security team would put them. In a cloud deployment they come from AWS Secrets Manager or Azure Key Vault, injected at start-up — the product detects that it is hosted and stops offering to store secrets locally at all. On a laptop, a key entered in the settings is sealed with the workspace key, kept at file permissions only you can read, and never shown again — only the last three characters, so you can tell which key is in use without exposing it. While you are typing it, an eye button shows what you have entered, because a mistyped key is otherwise a silent failure.
Is it learning from our data?
No. No customer content is used to train any model, in any deployment option. The model services used (in your own cloud, or via API) do not train on your content either. The tool also keeps a full log of every model call — what was sent, in what protected form, at what cost — so your own team can verify the claim rather than take it on trust.
How do our people sign in?
Cloud deployments sign in through your own directory — Entra ID or Amazon Cognito — with multi-factor authentication required without exception; an instance is not handed over until a test login has been MFA-challenged. The tool verifies the identity token itself rather than trusting whatever sits in front of it, and a group you nominate decides who may use that instance, so someone who leaves loses access when their account does. There is deliberately no separate application password beside it: a shared password that outlives a leaving date is the thing the directory was brought in to prevent. On a laptop, where there is no directory, a workspace password does the job instead. How each option is arranged.
Working with it
What do we need to provide?
A sponsor, about fifteen minutes of leadership interview, and a folder of documents the business already produces. The tool tells you what it can and cannot test with what you gave it, and produces a request list — with owners — for the evidence it is missing. Thin evidence is reported as a finding, not papered over. It also asks for your customer frame — who you serve, where those customers are heading, what they need, and who you will not serve — because no customer, no business: each statement becomes a testable commitment, checked against the customer evidence you already hold (CRM, win/loss, complaints, satisfaction data).
Who sees the findings?
The sponsor who commissioned the work, and the adviser running it, first. The executive team sees candidates before any board does, with the right to dispute the facts — disputes are recorded on the finding. Nothing reaches a board as a verdict; it arrives as evidence with judgement attached, and the humans decide.
Does it replace our advisers?
The opposite: it gives them better raw material. The tool compresses weeks of diagnostic reading into hours; the conversations that follow — what to do about a real contradiction, whether to change the behaviour or change the strategy — are precisely the work only a trusted adviser can do. Every quarterly run generates those conversations rather than replacing them.
How long does an engagement take, and what does it cost?
A testability review is two half-days plus pre-work. A baseline audit runs over one quarter's documents. The compounding value is the quarterly rhythm after that. Pilots are individually priced, with written exit criteria agreed before anything starts — you will know, in advance and in writing, what a successful pilot has to show you.
Can it help us formulate the strategy, not just check it?
The first piece of this is live: from your charter and your evidence, the tool derives next moves — where to cut, where to invest, and where a partnership shape fits — always as competing options with consequences, never as instructions. Every move is grounded in verbatim quotes from your own documents, a proposed cut states the capability it touches and the protective floor that should ship with it, partnership moves describe the shape required (never a named company), and each move waits for a named person's decision, which enters the record either way. The fuller formulation assistant — drafting the strategy itself as testable choices from your problems, opportunities and aspirations — is in development, designed with the same care, because "AI wrote our strategy" is exactly what this must never become.
Can it show where the strategy is under-resourced — people, technology, process?
Yes, with a discipline that keeps it honest: no claim, no light. A people, technology or process row lights only where a commitment your leadership actually confirmed leans on that kind of evidence — and then it speaks plainly: stop (a commitment in conflict, or its required evidence absent), pause (caveats or partial evidence), go (evidenced, nothing adverse). Where the evidence to draw a row does not exist, it says "not drawn" rather than showing a reassuring green — an honest grey is a finding, not a clean bill. And a red is never advice: the way out of it is a set of options with consequences, decided by a named person.
How would we know the strategy is actually being implemented?
This is where the roadmap is pointed. Each commitment carries an owner, the chain of people responsible beneath them, and — the part that matters — a distinction between three states: evidenced (a document shows it moving), asserted (someone said "on track", recorded as exactly that), and silent (nothing either way). "Seven of twelve commitments are progressing on evidence, three on assertion only, two are silent — and here is who each is waiting on" is a sentence a board can actually govern with.
When can we use it?
It is built and in private pilots now with a small number of boards, investors and advisers. Start with the five-minute readiness check — it tells you whether a review would have something real to test, and if not, exactly why — or register interest.
Still deciding?
Take the readiness check
The readiness check takes five minutes, is scored in your browser, and stores nothing unless you choose to send it. Either way, you learn something.