Methodology
This page is the public rule book behind everything on atozaiaccounting.com. It explains what the scores mean, where the data comes from, what a vendor can change, what only the editor can change, and what's on the roadmap that isn't built yet.
It is a working document. The site is updated several times a day; this document is updated less often. When the two disagree, the site wins and the next revision of this document catches up.
1. What this is, and what it is not
The A to Z of AI Accounting Software is two things:
- A live, independent directory of vendors operating in the AI accounting and finance stack.
- A seven-chapter editorial report on where the category is going and who is positioned to win it.
The directory is structured, filterable, and updated continuously. The report is editorial. The report does not change every time a vendor adds a new feature. It evolves with the market, not the press cycle.
I am not a vendor and I am not an analyst firm. I am a former practitioner writing what I see, with the bias that comes from having been on the buy side and the build side. I do not take pay-to-play money from vendors for placement, ranking, or commentary. Vendors can claim their entry for free, which gives them edit rights on factual fields. Vendors cannot pay for a higher AI Maturity score, a place on the Map, or favourable copy in Our Perspective.
There are mistakes in this directory. Some of them are mine, some of them are out-of-date public information, some of them are the vendor's own copy. I publish the rules I use so you can spot the mistakes faster than I can.
2. How updates land
The Directory is updated continuously. When a vendor's claim is approved or an edit request goes through, paid users see the change within seconds. The free preview catches up within the day, usually within the hour.
The Report is a different beast. Chapter content is editorial. It revisits material market shifts rather than tracking individual product launches. Expect material revisions every four to eight weeks, not daily.
Confidence scores (see §5.4) recompute every Monday morning. Map composition and Directory rankings recompute on every page load using the current data and the algorithm in §4.2.
3. The Stack: how vendors are categorised
Every vendor sits in a single Main category that describes where it does the work, plus a Sub-category that picks the specific job. Vendors that genuinely span two main categories (rare) get the one closest to their centre of gravity.
The ten main categories:
- Economic Event: where money moves or commits. Banks, payments, expense capture at source, anything upstream of the ledger.
- Ingestion: capture and extract source documents. OCR, document capture, supplier invoice ingestion.
- Normalisation: raw transactions into clean, reconciled, categorised data. Bank reconciliation, transaction coding, close.
- General Ledger: the authoritative ledger of record. Both AI-native ledgers and traditional ERPs sit here.
- Assurance: audit, internal audit, controls, SOX, risk and compliance.
- Tax: direct, indirect, personal, corporate compliance and filing.
- Decision & Action: FP&A, AP, AR, spend management. Anything that consumes the cleaned ledger to drive a decision or a payment.
- Orchestration: the control layer above the ledger. Practice management, agent management, workflow automation for firms.
- Productivity: horizontal tools repurposed for accounting (call transcription, document automation that isn't accounting-native).
- Data Layer: infrastructure under everything. APIs, identity, analytics pipes. Not user-facing.
Sub-categories are the working list. I add new ones when enough vendors arrive in the same place to justify a name; I rename them when the language settles. The current list is published in the Directory filter.
A vendor's category is editorial. Vendors can challenge it through the claim form. I read every challenge.
4. The Map: who's in and who's out
The Map shows the most relevant new players driving innovation across each centre of gravity. It is not "the biggest" and it is not "the best". It is an editorial pick with a transparent algorithm on top.
4.1 What gets a vendor onto the Map
Two gates:
- Editorial flag: I have to mark the vendor as in-scope for the Map. This is a human call. The bar is: "does this vendor materially shape the conversation in its category".
- Launch status: Live and Beta vendors are eligible by default. Alpha and Waitlist vendors are normally directory-only, not on the Map. They are visible in the Directory with a clear launch-status pill.
4.2 The Big 30 algorithm
Once eligibility is established, the Map shows up to 3 vendors per main category, except General Ledger which gets 6 because that's where the largest re-platforming bet sits.
Within each category, ranking is:
- AI Maturity descending (a 4 sits above a 3 sits above a 2).
- Verified descending (claimed vendors sit above unclaimed at the same maturity).
- UK presence descending (UK-operating vendors sit above non-UK at the same maturity and verification).
- Name ascending (alphabetical tie-break).
If a category has fewer eligible Live or Beta vendors than its quota, the algorithm tops up from the Alpha and Waitlist pool to fill the slots, ranked the same way. When that happens the launch-status pill on the card tells you the vendor is early.
4.3 Why some big names are not on the Map
Two reasons:
- They are not AI-native, and including them would dilute what the Map is for. The Directory is where you find every vendor; the Map is a curated read of where the centre is moving.
- The category quota is full and a smaller, more interesting vendor is doing more on AI right now. Quota matters more than headline size.
You can dispute either call through the claim form or by emailing me directly. I move vendors on and off the Map roughly weekly.
5. The scores: what each one means
5.1 AI Maturity (0–4)
This is the most important score on the site. It is also the one most worth disputing if you think it's wrong.
The rubric:
- 0 — No material AI. Traditional rule-based software. AI is not part of the product story.
- 1 — AI as feature. A chatbot, basic extraction, or narrative summary layered onto an otherwise unchanged product. Marketing emphasis is on AI; product reality is mostly the same product with an AI add-on.
- 2 — AI-assisted workflows. AI meaningfully speeds up or improves individual capabilities but humans remain the primary decision-makers across the product. Typically one or two capabilities affected.
- 3 — AI-led in core tasks. AI is the primary decision-maker in at least three of: categorisation, reconciliation, anomaly detection, compliance checking, forecasting, narrative generation, workflow routing. Humans handle exceptions and review.
- 4 — Agentic end-to-end. The product delegates whole workflows to AI agents with human oversight rather than human execution. The product is recognisably built around agents, not around screens.
AI Maturity is editorial. Vendors cannot edit it directly. They can submit a challenge through the claim form, with evidence (product demo, public roadmap, customer reference). I review challenges within a few business days.
This score is harder than the others to keep current. It is the area I am putting most automation work behind: news-watching for product launches, evidence ingestion from public sources, and a structured re-review cadence. Expect this rubric to harden in v2.
5.2 AI Core to Value
A three-way: Yes / Partial / No.
- Yes: remove the AI and the product either doesn't work or loses its core differentiation.
- Partial: AI is a meaningful part of the value but the underlying product also works without it.
- No: AI is a marketing wrapper or a tiny feature.
A vendor can have AI Core to Value = Yes and a maturity score below 4. Plenty of agentic products are at maturity 2 today and will be at 4 in two years. The two scores answer different questions: "is AI structurally what this is" vs "how far along is the AI today".
5.3 Verified ✓
A vendor is Verified when their claim has been approved by me. Claiming is free. The vendor confirms or corrects every public-facing field on their record, and the Verified badge appears.
What Verified means:
- A real human at the vendor has reviewed the entry.
- The factual fields (category, launch status, ownership, funding, modules, use cases, regions, and so on) are accepted by the vendor as accurate.
- Editorial fields (Our Perspective, AI Maturity, Threat to Incumbents) remain editorial. Verified does not mean "the vendor approves of the perspective written about them".
Verified is a public signal. Vendors are welcome to call out their Verified status off-site through the share kit on their company page; the badge tells readers a human at the company stood behind the factual entry, nothing more.
What Verified is not:
- It is not a quality endorsement. A Verified vendor at AI Maturity 0 is just a verified vendor that doesn't ship AI.
- It does not expire automatically today. If a vendor goes stale (12+ months without contact), the Confidence score downgrades and the badge is paired with a Medium or Low confidence indicator. A future revision will retire badges for vendors that haven't responded for two cycles.
5.4 Confidence / Information Integrity
A three-way: High / Medium / Low. Recomputed every Monday morning.
- High: Confirmed with the vendor in the last 12 months.
- Medium: Cross-checked across multiple public sources within the last 18 months, or vendor-confirmed but ageing past 12 months.
- Low: Public sources only, or stale, or thin (three or more of the six key fields blank: public AI use cases, workflow stages, core modules, main category, ledger requirement, AI maturity).
Confidence is computed, not editorial. It is the honest answer to "how much should you trust this entry today". An unclaimed Low-confidence entry is a flag to me as much as to you.
5.5 Threat to Incumbents
A three-way: Low / Medium / High. Editorial only.
- Low: complements rather than competes; niche; enabling layer; or an incumbent holding ground with no disruptive pressure of its own.
- Medium: competes meaningfully with existing players; capturing share in a contested segment; credible but not dominant.
- High: dominant in category, actively displacing incumbents, or posing a clear existential threat.
Vendors can challenge this through the claim form.
5.6 Launch Status
- Live: generally available, customers can buy and use the product today without restriction.
- Beta: available to limited customers (invite, signup approval, capacity-limited rollout).
- Alpha: early access only. Design partners, hand-picked customers, pre-release testing.
- Waitlist: not yet shipping. Collecting signups for a future launch.
Live and Beta are eligible for the Map. Alpha and Waitlist are Directory-only by default but can land on the Map if a category quota would otherwise be unfilled (see §4.2). Vendors can challenge this through the claim form.
5.7 MCP Depth & Safety (0–5)
Model Context Protocol, or MCP, is an open standard that lets an AI assistant like Claude or ChatGPT connect directly to a software product and work with your data. A vendor builds one MCP server and any assistant that speaks the protocol can use it.
I score how good that server is, not whether one exists. I deliberately do not count tools. Tool counts are easy to inflate, impossible to verify from outside, and vendors' own published figures frequently do not reconcile with each other. What I look at instead is what an accountant would actually want to know before connecting an agent to a live ledger: can it change your data, how narrowly can you fence in what it touches, and is there a record of what it did.
- 5: reads and writes, with a log of every AI action.
- 4: reads and writes, with fine-grained permissions.
- 3: reads and writes, with account-wide permissions.
- 2: reads your data, cannot change anything.
- 1: announced, no public documentation.
- None found: I checked and found no MCP server.
- Could not check: the vendor's site blocks automated checks.
The line between 3 and 4 is the one that matters most in practice. Both can change your books. A tier 3 server hands an agent everything the account can reach, so an assistant connected for one client can act across all of them. A tier 4 server can be restricted to a single entity, client or action. Same capability, very different consequences if something goes wrong. Tier 5 adds the record: scoping limits what an agent can damage, a log lets you find out what it did. Very few vendors have both.
What I check, and what the scores do not tell you. I read what a vendor publishes: their site, their developer and documentation subdomains, their llms.txt, and whether an MCP endpoint responds. Everything is dated, because this moves quickly. That means I am scoring published documentation, not testing a live server against real data. A vendor with an excellent MCP behind a customer login will score lower here than one with tidy public docs. I think the trade is worth making because it can be applied consistently to every vendor and re-run every month, but it is a real limit and I would rather state it than imply more than I know.
“None found” does not mean “no MCP.” It means I checked that vendor's own site on that date and found nothing. The method has been wrong before: one vendor was checked three separate times and recorded as having nothing, and a better check later found around fifty tools documented on their own site. Treat a positive finding as solid and an absence as provisional.
“Could not check” means the vendor's site blocks automated requests. That is ordinary enterprise security and not a mark against anyone. It is more common among large vendors than small ones, so this state is disproportionately the household names.
When it counts toward the ranking. MCP Depth & Safety is published on vendor profiles from the September 2026 edition and enters the World Ranking composite from October 2026, so vendors have notice before it moves anyone's position. Vendors with no MCP, and vendors I could not check, are treated neutrally rather than penalised. A good MCP can lift a score; not having one does not push anyone down.
6. Editable vs editorial: who controls what
A vendor can claim their entry and edit the factual record. The editorial layer stays with the editor.
6.1 Vendor-editable after claim
Through the claim form and subsequent edit requests, a vendor can update:
- Identity: name, website, CEO, ownership, founded year, funding stage, year, total raised band.
- Footprint: UK presence, operating regions, target market.
- Product: core problem solved, core modules, workflow stages covered, public AI use cases, ledger integrations, ecosystem connectivity, AI Core to Value.
- Commercial: primary user, primary buyer, end user, sales motion, pricing model, partner dependency, growth scale, requires core ledger.
- Category: main category, sub-category (challenge only — admin reviews).
6.2 Challenge-only (admin reviews with reasoning)
- AI Maturity
- Launch Status
- Threat to Incumbents
A vendor submits proposed value plus evidence; admin reviews and either updates the record or rejects with reason.
6.3 Editor-only, never vendor-editable
Limited to fields the public can see on the site:
- Our Perspective: the editorial paragraph on the company page.
- Verified ✓ badge: set by the editor on claim approval.
- Confidence score: computed weekly (see §5.4); not editor-set.
- Map inclusion: whether a vendor appears on the Map at all (see §4).
- MCP Depth & Safety: assessed from published documentation (see §5.7). A vendor can send the URL of their MCP server or its documentation and I re-run the check against it, the same way AI use-case evidence works. Vendors supply the link, the assessment stays mine.
A handful of further editorial signals are captured behind the scenes (working notes on threat rationale, internal categorisation, source trail, and so on). Those are not published today and don't sit on the public record.
7. The other classification fields
Brief descriptions; full enum lists live in the Directory filters.
- Workflow stages covered: the steps in an accounting workflow this product touches. Capture, categorisation, reconciliation, approvals, close, consolidation, reporting, variance & narrative, forecasting & planning, audit & assurance, plus adjacent stages like client onboarding, payroll, practice management, statutory & secretarial, tax.
- Core modules: the named features inside the product. A fifty-something-strong list ranging from Bank Feeds & Reconciliation to Agentic Workflow Builder.
- Public AI use cases: the AI capabilities the vendor publicly advertises. Used to validate AI Maturity claims and to power filtering. If the vendor lists capabilities but the product doesn't demonstrate them, the maturity score will not reflect the marketing.
- Requires core ledger: Yes / No / Partial / N/A. Whether the product needs a general ledger underneath it to function.
- Sales motion: Self-serve / Sales-led / PLG / Partner-led / Developer-led / Hybrid.
- Pricing model: Transparent / Tiered-public / Gated-demo / Enterprise-only.
- Ownership: Founder-led / VC-backed / PE-backed / Public / Acquired / Corporate subsidiary / Government.
- UK Presence: Yes / Limited / No / Coming soon. "Yes" means the vendor has real customers and named UK go-to-market motion. "Limited" means UK customers exist but the product is not localised or supported. "No" is no UK footprint today. "Coming soon" is a stated commitment.
8. Research methodology
Three inputs feed every entry:
- AI-assisted desk research from public sources. Initial drafts of every record are produced from web data, product sites, press releases, funding databases, and direct product documentation. AI is used to extract, normalise, and structure. AI does not make editorial calls.
- Primary research through interviews. I have detailed conversations with vendors covering product roadmap, demos, customer base, and AI architecture. Not all vendors respond. Where a vendor has responded, the entry is marked Verified and Confidence reflects recency.
- External views from customers, accountants, and investors I speak to. These rarely show up as a cited source but they materially inform Our Perspective. The Our Perspective paragraph is a composite read; it is not a quote from any single interview.
Where a vendor has not responded, the entry is built from public sources only. Confidence is set Low. Unclaimed entries are clearly marked as such on the site and in the Directory.
Sources are captured internally for every entry. They are not currently published per-vendor (which is a known gap; see §10).
9. Update cadence
- Vendor records: live for paid users, daily-or-better for free preview. Approved claims and edit requests land within seconds.
- Confidence scores: recomputed every Monday morning.
- Map composition and Directory rankings: refreshed on every page load using the current data and the algorithm in §4.2.
- Map inclusion: I review the editorial in/out call roughly weekly.
- Report chapters: editorial. They change when the market changes, not when a vendor ships a feature. Expect material revisions every four to eight weeks.
- This methodology document: versioned with the date stamp at the top. v2 lands when the AI Maturity rubric hardens, news-watching is automated, and source citations go public.
10. What's wrong with this today, and what I'm doing about it
Things I know are wrong or weak today:
- The Maturity rubric is still calibrating. Edges between 2 and 3 are the hardest call. Disputes are encouraged.
- Sources are not public per-vendor. They will be. The plumbing exists, the surface doesn't.
- Some vendors are out of date. The Confidence score is the honest answer to "how much should you trust this today". A Low-confidence entry is a flag.
- Verified does not yet auto-decay. If a Verified vendor goes 18 months without contact, the badge should retire. In v1 it doesn't. Confidence downgrades instead.
- Threat to Incumbents notes are not public. The score is; the rationale is editor-only today. I want this to be public in v2 so the call is reviewable.
Things I am building toward:
- News-watching per vendor: automated ingestion of product launches, funding, and personnel moves into the entry's source trail.
- Structured re-review cadence: every Verified entry gets a touch every 12 months minimum, automated reminders to the vendor.
- Public source trail: per-entry list of where each fact came from, with last-checked date.
- Maturity v2: harder rubric with evidence requirements per level, made auditable.
- Editorial changelog: per-vendor changelog of when Our Perspective was last revised and why.
11. The World Ranking and the Big 30: how the lists are built
There are two different "who's leading" surfaces on the site, and they answer different questions.
11.1 The Big 30 (on the Map)
The Big 30 answers: "who's the sharpest pick in this specific category, right now?" It's deliberately capped at 3 vendors per category (6 for General Ledger) so no single category swamps the Map, and it's sorted the simple way described in §4.2: AI Maturity first, Verified status second, UK presence third, name as the final tie-break. It rewards depth in a lane, not overall size.
11.2 The World Ranking
The World Ranking (published monthly, top 50) answers a different question: "who's furthest along on real AI capability, full stop, regardless of category?" Every vendor in the report gets a single 0–100 score and is ranked against every other vendor directly, not just within its own category.
The score is built from five ingredients. In the interest of keeping the list resistant to gaming, I publish what goes in, not the exact weighting of each part:
- How mature and central the AI actually is — the AI Maturity rating, plus whether AI is core, partial, or cosmetic to the product (see §5.1 and §5.2).
- How advanced the claimed capabilities are — foundational capabilities (document extraction, categorisation) count for meaningfully less than advanced ones (agentic workflows, autonomous close, tax intelligence). A claim that hasn't been through evidence review yet still counts, but at a discount versus one that has.
- How much of the workflow the product actually covers — breadth across modules and workflow stages (§7).
- How well-evidenced and how visible the listing is — claimed status, Confidence (§5.4), genuine press coverage, and how deeply the vendor is covered in the report itself.
- A structural discount for AI that isn't core to the product — a large, well-known incumbent with a bolted-on AI feature is discounted on the size/coverage-driven parts of the score, so it can't out-rank a smaller AI-native challenger purely by being bigger or better known.
The World Ranking is a general capability ranking. It is deliberately not personalised, and it is not a recommendation for your specific business — for that, see My Stack (§12) or use Ask. Movement (▲/▼) shows change since the previous published edition, once more than one edition exists to compare against.
12. My Stack (the Hub): personalised, not general
Everything above — the Directory, the Map, the Big 30, the World Ranking — is the same for everyone. My Stack is the opposite: it's personalised to you, available to Full Access subscribers, and it's where "what's right for me" actually gets answered rather than "who's leading overall".
My Stack has four views:
- Recommendation: a shortlist filtered to your own profile and the problem you're trying to solve, not a general leaderboard position.
- Saved: track the vendors you're actually using, evaluating, or keeping on the radar — your own working stack, not the site's.
- Watching: live signal (funding, launches, announcements) for vendors you've pinned, so you don't have to keep checking back.
- Compare: side-by-side comparison of your saved vendors, built for the shortlist you're actually deciding between.
The framing adapts to who you are (practitioner, analyst/investor, or other) but the underlying data is the same. If the World Ranking and Big 30 tell you who's leading in general, My Stack is where that gets narrowed down to who's leading for you.
13. How to dispute or correct an entry
Three routes:
- Claim the entry. Use the "Claim this entry" link on any unclaimed company page. Free. You get edit rights on factual fields and a chance to challenge editorial fields with evidence.
- Edit request. If you've already claimed, hit "Edit this entry" on your company page. Goes into the editorial queue.
- Email. hello@atozaiaccounting.com. For anything that doesn't fit the structured forms (a missing vendor, a wrong category, a perspective you think is materially off).
I read everything. I do not always agree.
Versions
- v1.1 · 10 July 2026: added §11 (World Ranking and Big 30 list-building) and §12 (My Stack). What's published is what goes into the World Ranking score, not the exact weighting — see §11.2.
- v1.0 · 17 May 2026: first published methodology. Covers everything above.