Autonomous · Hierarchical · Governed · Verifiable

Give it a mission. Get back the work — and the proof.

Butai is SHINRAI’s agentic AI workforce. It plans the mission, composes a crew of world-expert agents, executes through typed handoffs, gates its own quality, and records every decision in a hash-chained ledger — then seals the result in an evidence pack you can verify, all inside budgets that guarantee it stops.

  • Planned & QA-gated
  • Human-in-the-loop when it matters
  • A signed, verifiable evidence pack
mission_trace · rate-limiter · design + benchmark + reviewSample
seqeventactorrun
01planobjective + definition of donenode:plan7a1c…e04
02composeengineering + research + QAnode:composer7a1c…e04
03route→ engineeringchief:supervisor7a1c…e04
04artifactspec publishedeng:architect7a1c…e04
05tool_callcode · sandboxedeng:backend7a1c…e04
06clarify◆ paused for a human answerhuman:operator7a1c…e04
07handoffeng → researchchief:supervisor7a1c…e04
08artifactbenchmark analysisres:data-scientist7a1c…e04
09qa_gate◆ verified vs definition of doneops:qa-reviewer7a1c…e04
10finalizeanswer synthesised chief:supervisor7a1c…e04
Every routing decision, tool call, artifact and verdict is recorded in a hash-chained ledger. ◆ marks a QA gate or a human-in-the-loop pause.chain verifiable
  • Hierarchical teams of world-expert agents
  • Signed, verifiable evidence packs
  • Provider & region failover
  • Nairobi · Dubai

Why Butai exists

One agent is easy. A workforce you can trust is not.

A single prompt can draft something plausible. Delivering a real objective takes a team that plans, divides the work, checks itself, and leaves a record of how it got there.

Who did what?Work is routed by a Chief Supervisor to teams, and by each Team Supervisor to named expert agents. Every routing decision and tool call is attributed and recorded.
How do they collaborate?Experts publish typed work artifacts to a shared blackboard and hand off explicitly — within and across teams — instead of relying on raw chat context.
When is it done?Planning fixes a checkable definition of done. A QA gate verifies the work against it; only an approved result finishes, and revisions loop back a bounded number of times.
Will it run away?Every mission runs under a model registry, a tool gateway, an autonomy ladder, and hard LLM-call and wall-clock budgets, so every run terminates with a controlled blast radius.
Can you prove it?The run is recorded in a hash-chained, tamper-evident ledger and sealed in an evidence pack — with a per-artifact content manifest — that can be signed so a third party verifies it with a public key alone.
What if a model falls over?A transient fault retries; a sustained outage trips a circuit breaker and fails over to the next healthy region, preserving governance. If every endpoint is down the run is held and resumes when one recovers.
Nobody marks their own work?The acceptance gate is a separate QA reviewer, and the independent reviewer is run on a different model family than the builder — a stronger check than role separation alone.
Does it know when it’s stuck?Beyond budgets, a no-progress detector stops a run that keeps failing the same way, and a per-run turn cap bounds pathological routing — so a run can’t spin without end.

How a mission runs

Plan, compose, execute, gate, finalize.

  autonomous agents, inside fixed budgets  a point where a human can decide

Every mission follows the same path through a hierarchical LangGraph. Teams of world-expert agents do the work; a QA gate decides when it is finished; and the whole run is bounded so it always terminates.

  1. 01planPlanner

    Sharpens the mission into an objective and a checkable definition of done, and lists the teams and steps it will take.

  2. 02composeComposer

    Assembles the crew for this mission — from the plan or by relevance — and always includes the QA gate team.

  3. 03routeChief Supervisor

    Routes the mission among the composed teams, hop by hop, until the work is ready for review.

  4. 04executeTeam Supervisor · Experts

    Each team supervisor routes to its world-expert members, who use their granted tools and publish typed artifacts to the shared blackboard.

  5. 05humanclarifyOperator

    If the mission is genuinely ambiguous, the run can pause mid-flight on a real interrupt, ask a person, and resume with the answer injected.

  6. 06handoffAny team

    Experts hand off explicitly — within a team or across teams — carrying the artifacts the next recipient needs. No shared context by accident.

  7. 07gateqa gateQA Reviewer

    The Operations QA Reviewer verifies the work against the definition of done. Only an approved result finishes; a revision loops back, bounded.

  8. 08finalizeChief Supervisor

    The Chief synthesises the polished final answer from the approved artifacts and the mission closes.

Revisions have limits.A QA verdict of needs-revision routes back to the Chief with the finding. The loop is bounded by a configurable maximum before the run is forced to finish.
Budgets guarantee it stops.A hard cap on LLM calls and a wall-clock budget are checked at every hop, so no mission can run unbounded or escalate cost.
A pause is not a failure.A run awaiting a human answer is held, not killed; a durable checkpoint lets it survive a restart and resume exactly where it paused.

Governance

Autonomy with the guardrails built in.

Six enforcement seams sit under every mission. Each one is a boundary an agent cannot talk its way past. Choose a seam to see what it enforces and how.

Model registry

No agent ever references a raw provider model id. Every LLM resolves a logical alias through a governed registry that enforces policy on the way through.

What it enforcesHow
Only approved models are usedAliases like expert-default resolve to a model with an approval status; unapproved models are refused at resolution
Data-privacy tier is respectedEach alias carries a data-tier and guardrail metadata the enforcement path checks before returning a client
One governed path for every backendLocal manifest, AWS SageMaker Model Package Groups and Snowflake Cortex adapters share a single enforcement path — none can bypass it
Every resolution is auditableAlias resolutions are logged, so you can see which governed model served which mission

These seams are enforced in code, not by policy documents. They are configurable per deployment and surfaced, with every admin action, in the operator console’s audit log.

Decision trace & evidence

Every mission leaves proof of how, not just what.

The trace is built from the run itself, not written from memory. It records every routing decision and its rationale, every gateway-checked tool call, every typed artifact published to the blackboard and the handoffs that carried them, the QA verdicts and revision loops, any plan deviations, and the human clarifications folded back in — written to a hash-chained, tamper-evident ledger a verify endpoint can check.

At the end, the run is sealed in an evidence pack: an engine-computed per-artifact SHA-256 manifest, the subprocessors routed through, an autonomy declaration, per-section AI-authorship disclosure, and the unit economics. It can be signed with Ed25519 so a third party verifies it with a published public key alone — no shared secret, no trust in our tooling.

SurfaceWhereWhat it shows
Live streamSSE on the run endpointplan, chief routing, member activity, artifacts, QA verdict and the final answer, as they happen
Admin portalThe /admin consoleEvery run’s status, current team · agent, budget usage, the full artifact + handoff trail, and metrics
Hash-chained ledger/trace/verifyAn append-only, canonical-JSON record of routing, tool calls, verdicts and deviations — verifiable, first broken link reported
Evidence pack/evidence-pack + /evidence/public-keyJSON or client-readable Markdown with the artifact manifest, provenance and an optional third-party-verifiable signature

A trace shows how a result was reached and what it cost. The evidence pack lets someone else confirm it — without taking our word for it.

How the work holds together

Structure instead of a long chat.

Collaboration is explicit. Experts don’t share a growing transcript — they publish typed artifacts and hand off deliberately, so every recipient gets clean, deterministic context.

two-level supervision

A Chief routes to teams; each team routes to experts.

A Chief Supervisor picks the next team. Each Team Supervisor picks the next world-expert member. Routing is a structured choice among the allowed options, not a free-for-all.

world-expert members

Specialists, not a single generalist.

Engineering spans the full SDLC — product, spec, architecture, backend, frontend, data, SDET, QA, security, DevOps, docs — alongside Research, Content, Customer and Operations teams.

typed artifacts

Work products, not messages.

Each expert declares the artifact types it produces — requirements, spec, design, code, test plan, review, analysis — and publishes versioned artifacts to a shared blackboard.

structured handoffs

Deliberate transfers.

Handoffs name the sender, recipient, intent and the artifacts passed — intra-team, cross-team, escalation or return — so context moves on purpose, not by accident.

QA acceptance gate

Nobody marks their own work.

A dedicated Operations QA Reviewer verifies the mission against its definition of done. Only an approved result finishes; a needs-revision verdict routes back, bounded.

human-in-the-loop

A real pause, not a guess.

When a mission is genuinely ambiguous, the run interrupts and checkpoints, an operator answers, and it resumes with the answer injected — surviving a restart if needed.

Models & tools

Run it anywhere. Keep one governed path.

Agents resolve a logical model alias through a governed registry, and reach for tools through one gateway. Swap the backend to fit your stack — the enforcement path, and the agents, don’t change.

Registry backendResolves aliases toGovernance read fromNeeds a cloud?
Local manifestModels in a YAML/JSON manifestApproval status, data tier and guardrails in the manifestNo — runs fully local
AWS SageMakerModel Package Groups (latest Approved package)Package approval status and governance metadataAWS account
Snowflake CortexModels in the Snowflake Model Registry, served via CortexVersion metadata on the registered modelSnowflake account

Teams & their tools

Tools are granted per team by least privilege, and every call is risk-classified by the gateway. Web search backs any team that needs current facts; a sandboxed Python tool runs untrusted code in a restricted subprocess; a read-only SQL tool is SELECT-only, allowlisted and row-limited against a replica.

Engineering · web · codeResearch & Analytics · web · code · SQLContent & Marketing · webCustomer & Sales · webOperations, Legal & Compliance · web · SQLQA gate

Deploy

Ship it with Docker Compose (Postgres + app, migrations applied on start), run it locally with an in-memory checkpointer and auth off, or drive it from the shinrai CLI. A Postgres-backed LangGraph checkpointer makes every conversation a durable, resumable thread.

Python 3.11LangGraphFastAPI + SSEPostgreSQLOAuth2 + JWTDocker

Missions & proposals

Two ways in. One governed workforce.

Run a mission directly on a thread, or go through the governed proposal lifecycle — scope a raw request into a costed proposal, accept it, deliver it, and revise it with feedback. Both paths run the same instrumented graph.

The governed path

The proposal lifecycle

A raw client request is scoped at intake into a costed proposal an operator can accept or reject. Accepting and delivering dispatches the full workforce on a dedicated, tracked thread. A delivered proposal can be revised with feedback, which dispatches a fresh run honouring the original scope plus the change.

A · IntakeB · DeliveryC · Revision
proposal statusLifecycle
intake bounded scoping pass
Scopedproposednot_readyclarifying
Decisionacceptedrejected
Deliverydeliveredrun linked
Revisionfeedbackfresh run

scoping and delivery run the workforce graph accept / reject / deliver by an operator

Pipeline A — Intake

  • A deliberately cheap, bounded scoping pass: readiness gate → plan → deliverables, metered for an estimated cost
  • A ready request becomes a proposed proposal with objective, deliverables and plan
  • A not-ready one becomes clarifying with an auto-generated clarification interview

Pipeline B — Delivery

  • Accepting and delivering dispatches the full workforce run through the run manager
  • Runs on a dedicated thread; the proposal moves to delivered and links its delivery run
  • Tracked as a cancellable task with its own per-run budget

Pipeline C — Revision

  • A delivered proposal can be revised with client feedback
  • Dispatches a fresh delivery run honouring the original scope plus the feedback
  • Every revision is linked back to the proposal it came from

Direct missions

  • Skip the proposal and run a mission straight on a thread
  • Autonomous end to end: plan → compose → execute → QA gate → finalize
  • Stream activity over SSE, or get the final answer in one call
Bounded

Hard LLM-call and wall-clock budgets per mission guarantee every run terminates.

Resumable

A Postgres checkpointer makes every thread durable — cancel, resume or restart a run.

Costed

Token accounting and an estimated USD cost from a governed price table, per run.

Multi-tenant

Threads are tenant-owned; checkpointer thread ids are namespaced per tenant.

The clarification interview

When scoping finds gaps, the platform raises a round of role-addressed questions — business, product owner, CTO, CISO, data protection, legal, finance, operations, end user — instead of guessing. Answers can be captured in the admin UI or exported as a Markdown, CSV or print-ready HTML questionnaire and re-imported. Every answer is attributed to the authenticated responder, and intake-origin rounds fold the answers back in and re-scope, bounded by a maximum number of rounds.

Role-addressedExport / importAttributable answers

The teams on the roster

A mission is staffed from a declarative org chart. Adding a team or an expert is pure data — the orchestration engine needs no changes.

web · code

Engineering

Full SDLC: product, systems analysis, stack strategy, UI/UX, architecture, backend, frontend, data engineering, SDET, QA, security, DevOps and docs.

web · code · SQL

Research & Analytics

Head of research, research analyst, data scientist, BI analyst and an insights synthesiser.

web

Content & Marketing

Content strategist, copywriter, editor and SEO specialist.

web

Customer & Sales

Support engineer, customer success manager and enterprise account executive.

web · SQL

Operations, Legal & Compliance

Compliance reviewer, operations analyst and the QA Reviewer that gates acceptance.

declarative

Your own teams

Declare a new team or expert — its remit, tool grants and artifact types — and the composer can staff it without touching the graph.

How Butai compares

A workforce, not a bigger prompt.

Assistants and frameworks help one model do more. Butai is a governed organisation of agents that plans, divides work, gates its own quality, and records how it got there.

QuestionSingle-agent assistantsAgent frameworksDIY orchestrationTypical AI consultancyBUTAI By SHINRAI TECHNOLOGIES LIMITED
Plans the mission & a definition of done◐Ad hoc◐You build it◐You build it◐In a document●Built in: objective + checkable definition of done
Hierarchical teams of specialists○One generalist◐Primitives only◐Hand-wired○Human-only●Chief → teams → world-expert members
Checks itself before finishing○No◐If you add it◐If you add it◐Varies●A QA gate others cannot skip
Governed models & tools○Raw model id○Your responsibility○Your responsibility◐Policy on paper●Model registry + risk-classified tool gateway
Records how the result was reached○Chat log◐Traces if wired◐If you build it◐Written afterwards●Hash-chained ledger + signed, verifiable evidence pack
Survives a provider outage○Call fails◐Retry if wired◐If you build it◐Manual●Retry, region failover, then hold & auto-resume
Guaranteed to terminate◐Context limit○Can loop○Can loop●Human-paced●Budgets, turn cap & a no-progress detector

A general comparison of approaches, not specific products. Capabilities of any tool vary and change over time.

Operate it

Dispatched, monitored, and under your control.

  1. 01 · Dispatch

    Every mission is a tracked run

    The run manager compiles a per-run graph with its own budget, trace recorder and cost accumulator, and tracks it as a cancellable task that heartbeats to the runs table.

  2. ◆ operator console02 · Monitor

    Down to the individual agent

    Watch each run’s status, current node and team · agent, step and call counts, budget usage, and the full artifact and handoff trail. A stale heartbeat flags a run as stuck.

  3. ◆ human decision03 · Control

    Cancel, resume, restart, clarify

    Resume continues the same thread from its last checkpoint, restart re-runs from the start, and clarify answers a run paused mid-mission and resumes it with the answer injected.

Who stands behind it

SHINRAI is the firm behind the platform.

Butai is SHINRAI’s own agentic workforce platform, and SHINRAI stands behind it. We are an AWS Advanced Tier Services Partner with offices in Nairobi and Dubai, delivering cloud, data and AI work for clients across industries.

AWS Advanced Tier Services PartnerSnowflake Select PartnerInformatica IDMC Partner
Founder & CEO

Timothy Munyao

Holds 13 AWS certifications and the AWS Golden Jacket, the first Kenyan to receive it. Leads Butai’s design and governance.

30%lower costs, about KES 1M saved a yearUnity Homes
40%fewer manual IT tasks, 20 hours saved a weekAdrian Group
83%faster deploymentsPallax Kenya
99.99%uptimePM Realty

Results from SHINRAI client engagements, as published on shinraitechnologies.io.

Fit

Who Butai is for.

A good fit

  • Teams who want multi-step work delivered by a crew of specialist agents, not a single chat
  • Organisations that need to see how an AI result was reached, and what it cost
  • Builders who want autonomy with a QA gate, budgets and governance already wired in
  • Multi-tenant deployments that need per-tenant isolation and auth from day one

Not a fit

  • A quick one-off answer where a single assistant is already enough
  • Fully unattended autonomy with no QA gate, budgets or human-in-the-loop
  • Work with no checkable definition of done anyone can agree on
  • Use cases that need a raw model id with no governance in the path

Questions

What teams ask first.

What is Butai, in one line?

An autonomous, hierarchical multi-agent platform. You give it a mission; it plans the objective and a definition of done, composes a crew of world-expert agents, executes through typed handoffs, gates its own quality, and finalises an answer — all recorded in a decision trace and bounded by budgets.

Is it fully autonomous?

It is autonomous within guardrails. An autonomy ladder (L0–L3) clamps how much it may do on its own, a QA gate decides when work is done, and budgets guarantee termination. When a mission is genuinely ambiguous, a run can pause mid-flight and ask a human, then resume with the answer injected.

How do the agents actually collaborate?

Through structure, not a growing chat. A Chief Supervisor routes among teams; each Team Supervisor routes to its experts. Experts publish typed, versioned artifacts to a shared blackboard and hand off explicitly, so each recipient gets clean, deterministic context.

Which models does it use?

Whichever your registry approves. Agents never reference a raw model id — they resolve a logical alias through a governed registry that enforces approval status, data tier and guardrails. Backends include a local manifest, AWS SageMaker and Snowflake Cortex, all sharing one enforcement path.

What stops it from running away or getting expensive?

Every mission runs under a hard cap on LLM calls and a wall-clock budget, checked at every hop. Each run gets its own budget so one stuck mission can’t drain the platform, a per-run turn cap and a no-progress detector stop a run that keeps failing the same way, and an optional hard deadline marks a run timed-out while leaving a checkpoint to resume from.

What’s the difference between a mission and a proposal?

A mission runs the workforce directly on a thread. A proposal is the governed path: a raw request is scoped at intake into a costed proposal an operator accepts or rejects, delivers as a tracked run, and can revise with feedback. Both run the same graph.

Can a human step in during a run?

Yes. On a genuine ambiguity the run raises a real interrupt that checkpoints the graph and holds as awaiting-clarification — not active, not terminal. An operator answers from the admin portal and the run resumes exactly where it paused, surviving a restart if needed.

Can we see how a result was produced — and prove it later?

Yes. Every mission writes a hash-chained, tamper-evident decision ledger — routing and rationale, tool calls, artifacts, QA verdicts, plan deviations and clarifications — plus token and cost totals, verifiable by an endpoint that reports the first broken link. The result is sealed in an evidence pack with an engine-computed per-artifact SHA-256 manifest, the subprocessors used, an autonomy declaration and AI-authorship disclosure. It can be signed with Ed25519 so a third party verifies it with a published public key alone (opt-in; an HMAC integrity signature is the default). The pack never asserts correctness or compliance — only what was done, by whom, and that the record is intact.

What happens if a model provider has an outage?

A single transient fault (dropped connection, timeout, throttling, 5xx) retries with bounded backoff. If an endpoint stays unhealthy, a per-endpoint circuit breaker trips and the call fails over to the next healthy region — rebuilt through the same governed path so approval, data-tier and guardrails are preserved. Same-provider region failover is on by default; cross-provider is opt-in and only ever uses enabled providers. If every candidate is down, the run is held (checkpointed, completed work preserved), the operator is notified, and it resumes automatically when an endpoint recovers or manually from the admin portal.

Can we add our own teams or tools?

Yes. The org chart is declarative data: a team declares its members, their tool grants and the artifact types they produce. Adding a team or expert needs no change to the orchestration engine, and new tools are registered in the gateway, which fails closed on anything unregistered.

Next step

Request a demo.

Tell us about the work you want a workforce to run. We’ll reply to arrange a 30-minute walkthrough on a mission of your choosing.

  1. Walkthrough. We run a mission end to end — plan, crew, execution, QA gate, trace.
  2. Your use case. We map your work to teams, tools and the registry backend that fits.
  3. Pilot. A scoped pilot on your own missions, with the decision trace to show for it.