North Stack vs generic AI agents

A browser agent gives you a transcript. We give you a ledger.

Operator, Claude computer-use and Lindy can click around your systems. What they can’t do is prove what they did or why, and they’ll take a different path on run two. In a regulated back-office, that’s the whole ballgame.

TL;DR

Generic AI agents improvise; North Stack runs a process you mapped. The agent decides on the fly what to do, the opposite of a pre-approved workflow, and leaves you a chat transcript rather than a tamper-evident record. North Stack triages the inbox into pre-mapped workflows, stops at human gates on anything sensitive, cites the exact knowledge-sheet row behind each decision, and writes every step to a hash-chained ledger you can hand to the FCA. If you need one-off personal research, a generic agent is the better tool, and we’ll say so below.

At a glance

The buying criteria that matter in a regulated back-office.

CapabilityNorth StackGeneric AI agent (Operator, Claude computer-use, Lindy)
Pre-mapped workflowYes: runs a process you approvedNo: improvises a path each run
Deterministic run-to-runYes: same map every timeNo: the same prompt can go two ways
Human approval gates on writesNative: hard gates on the sensitiveNone built in
Audit record typeHash-chained tamper-evident ledgerChat transcript
Decisions cited to a sourceRow-by-row to a versioned sheetBlack-box reasoning
Runs legacy UI, no APIYes: drives Acturis as-is, supervisedSometimes, but fragile, no gates
Team / governance modelRole-scoped access, approval chainsSingle power-user, no chain
Built forBack-office system-of-record workPersonal productivity & research

Determinism: a mapped workflow, not an improvised run

A general-purpose agent is non-deterministic and un-mapped by design. The same request can take a different path run-to-run because the agent decides on the fly what to do next. That is exactly the opposite of a pre-approved process. North Stack runs a workflow you mapped, with the exact steps, the systems each touches, and where the human gates sit, so it’s knowable before it ever runs, and every run is measured against the same map.

The audit record: hash-chained ledger vs chat transcript

An agent’s record of what happened is a transcript: not hash-chained, not tied to a versioned knowledge source, not structured for a regulator. North Stack appends every action, whether agent, human or system, to a ledger where each entry’s SHA-256 covers the one before it. Alter anything, anywhere, and verification pinpoints the first broken link. One click verifies the whole chain; one click exports it for a thematic review.

Human approval gates, or the absence of them

Reviewers of Operator specifically flag no approval chains and no governance model. North Stack makes the gate a first-class object: anything sensitive or uncertain stops at a person, with the proposed action laid out in full. Payments, bank-detail changes and regulator submissions are permanent hard gates that never auto-commit.

Cited decisions vs black-box reasoning

When a North Stack run decides, it cites the sheet, row and version it relied on, so you can trace the decision back to the rule behind it instead of believing a conversation. “The model decided” is not a defensible answer under Consumer Duty; “knowledge-sheet row 14, v3, said so” is.

Driving legacy systems safely

Browser agents can drive legacy UIs, but they hit login walls, CAPTCHAs and anti-bot blocking, and they do it with no gates and no role-scoping. North Stack operates the same Acturis screens a handler uses, under supervised, role-scoped access, resilient to screen changes, with the guardrails in the runner.

Who each is for

Two different jobs.

Choose North Stack when

You’re running regulated back-office work on a closed system of record (Acturis, insurer portals), you have a shared ops inbox with repeatable request types, a named compliance owner, and you need an audit trail you could hand to the FCA rather than a transcript you have to defend.

Choose a generic agent when

You want a single power-user’s personal productivity boost, like ad-hoc research, one-off browsing or exploratory tasks, with no compliance requirement and no approval chain. Operator and Lindy are better at that, and cheaper for it.