Autonomous · Hierarchical · Governed · Multi-tenant

Give it a mission. Get back the work — and the trace.

Butai is SHINRAI’s agentic AI workforce. It plans the mission, composes a crew of world-expert agents, executes through typed handoffs, gates its own quality, and records every decision in a ledger you can read — all inside budgets that guarantee it stops.

  • Planned & QA-gated
  • Human-in-the-loop when it matters
  • A queryable decision trace
mission_trace · rate-limiter · design + benchmark + reviewSample
seqeventactorrun
01planobjective + definition of donenode:plan7a1c…e04
02composeengineering + research + QAnode:composer7a1c…e04
03route→ engineeringchief:supervisor7a1c…e04
04artifactspec publishedeng:architect7a1c…e04
05tool_callcode · sandboxedeng:backend7a1c…e04
06clarify◆ paused for a human answerhuman:operator7a1c…e04
07handoffeng → researchchief:supervisor7a1c…e04
08artifactbenchmark analysisres:data-scientist7a1c…e04
09qa_gate◆ verified vs definition of doneops:qa-reviewer7a1c…e04
10finalizeanswer synthesised chief:supervisor7a1c…e04
Every routing decision, tool call, artifact and verdict is recorded. ◆ marks a QA gate or a human-in-the-loop pause.within budget
  • Hierarchical teams of world-expert agents
  • Proposals & missions
  • Postgres-backed, resumable runs
  • Nairobi · Dubai

Why Butai exists

One agent is easy. A workforce you can trust is not.

A single prompt can draft something plausible. Delivering a real objective takes a team that plans, divides the work, checks itself, and leaves a record of how it got there.

Who did what?Work is routed by a Chief Supervisor to teams, and by each Team Supervisor to named expert agents. Every routing decision and tool call is attributed and recorded.
How do they collaborate?Experts publish typed work artifacts to a shared blackboard and hand off explicitly — within and across teams — instead of relying on raw chat context.
When is it done?Planning fixes a checkable definition of done. A QA gate verifies the work against it; only an approved result finishes, and revisions loop back a bounded number of times.
Will it run away?Every mission runs under a model registry, a tool gateway, an autonomy ladder, and hard LLM-call and wall-clock budgets, so every run terminates with a controlled blast radius.

How a mission runs

Plan, compose, execute, gate, finalize.

  autonomous agents, inside fixed budgets  a point where a human can decide

Every mission follows the same path through a hierarchical LangGraph. Teams of world-expert agents do the work; a QA gate decides when it is finished; and the whole run is bounded so it always terminates.

  1. 01planPlanner

    Sharpens the mission into an objective and a checkable definition of done, and lists the teams and steps it will take.

  2. 02composeComposer

    Assembles the crew for this mission — from the plan or by relevance — and always includes the QA gate team.

  3. 03routeChief Supervisor

    Routes the mission among the composed teams, hop by hop, until the work is ready for review.

  4. 04executeTeam Supervisor · Experts

    Each team supervisor routes to its world-expert members, who use their granted tools and publish typed artifacts to the shared blackboard.

  5. 05humanclarifyOperator

    If the mission is genuinely ambiguous, the run can pause mid-flight on a real interrupt, ask a person, and resume with the answer injected.

  6. 06handoffAny team

    Experts hand off explicitly — within a team or across teams — carrying the artifacts the next recipient needs. No shared context by accident.

  7. 07gateqa gateQA Reviewer

    The Operations QA Reviewer verifies the work against the definition of done. Only an approved result finishes; a revision loops back, bounded.

  8. 08finalizeChief Supervisor

    The Chief synthesises the polished final answer from the approved artifacts and the mission closes.

Revisions have limits.A QA verdict of needs-revision routes back to the Chief with the finding. The loop is bounded by a configurable maximum before the run is forced to finish.
Budgets guarantee it stops.A hard cap on LLM calls and a wall-clock budget are checked at every hop, so no mission can run unbounded or escalate cost.
A pause is not a failure.A run awaiting a human answer is held, not killed; a durable checkpoint lets it survive a restart and resume exactly where it paused.

Governance

Autonomy with the guardrails built in.

Six enforcement seams sit under every mission. Each one is a boundary an agent cannot talk its way past. Choose a seam to see what it enforces and how.

Model registry

No agent ever references a raw provider model id. Every LLM resolves a logical alias through a governed registry that enforces policy on the way through.

What it enforcesHow
Only approved models are usedAliases like expert-default resolve to a model with an approval status; unapproved models are refused at resolution
Data-privacy tier is respectedEach alias carries a data-tier and guardrail metadata the enforcement path checks before returning a client
One governed path for every backendLocal manifest, AWS SageMaker Model Package Groups and Snowflake Cortex adapters share a single enforcement path — none can bypass it
Every resolution is auditableAlias resolutions are logged, so you can see which governed model served which mission

These seams are enforced in code, not by policy documents. They are configurable per deployment and surfaced, with every admin action, in the operator console’s audit log.

Decision trace

Every mission leaves a record of how, not just what.

The trace is built from the run itself, not written from memory. It records every routing decision and its rationale, every gateway-checked tool call, every typed artifact published to the blackboard and the handoffs that carried them, the QA verdicts and revision loops, any plan deviations, and the human clarifications folded back in.

It also carries the unit economics: token accounting and an estimated cost from a governed price table, totalled per run. An operator reads all of it in the admin portal — down to the individual agent — or over the API.

SurfaceWhereWhat it shows
Live streamSSE on the run endpointplan, chief routing, member activity, artifacts, QA verdict and the final answer, as they happen
Admin portalThe /admin consoleEvery run’s status, current team · agent, budget usage, the full artifact + handoff trail, and metrics
Decision ledgerThe trace storeA queryable record of routing, tool calls, verdicts, deviations and clarifications for audit

A trace shows how a result was reached and what it cost. It is the difference between an agent you hope worked and a workforce you can inspect.

How the work holds together

Structure instead of a long chat.

Collaboration is explicit. Experts don’t share a growing transcript — they publish typed artifacts and hand off deliberately, so every recipient gets clean, deterministic context.

two-level supervision

A Chief routes to teams; each team routes to experts.

A Chief Supervisor picks the next team. Each Team Supervisor picks the next world-expert member. Routing is a structured choice among the allowed options, not a free-for-all.

world-expert members

Specialists, not a single generalist.

Engineering spans the full SDLC — product, spec, architecture, backend, frontend, data, SDET, QA, security, DevOps, docs — alongside Research, Content, Customer and Operations teams.

typed artifacts

Work products, not messages.

Each expert declares the artifact types it produces — requirements, spec, design, code, test plan, review, analysis — and publishes versioned artifacts to a shared blackboard.

structured handoffs

Deliberate transfers.

Handoffs name the sender, recipient, intent and the artifacts passed — intra-team, cross-team, escalation or return — so context moves on purpose, not by accident.

QA acceptance gate

Nobody marks their own work.

A dedicated Operations QA Reviewer verifies the mission against its definition of done. Only an approved result finishes; a needs-revision verdict routes back, bounded.

human-in-the-loop

A real pause, not a guess.

When a mission is genuinely ambiguous, the run interrupts and checkpoints, an operator answers, and it resumes with the answer injected — surviving a restart if needed.

Models & tools

Run it anywhere. Keep one governed path.

Agents resolve a logical model alias through a governed registry, and reach for tools through one gateway. Swap the backend to fit your stack — the enforcement path, and the agents, don’t change.

Registry backendResolves aliases toGovernance read fromNeeds a cloud?
Local manifestModels in a YAML/JSON manifestApproval status, data tier and guardrails in the manifestNo — runs fully local
AWS SageMakerModel Package Groups (latest Approved package)Package approval status and governance metadataAWS account
Snowflake CortexModels in the Snowflake Model Registry, served via CortexVersion metadata on the registered modelSnowflake account

Teams & their tools

Tools are granted per team by least privilege, and every call is risk-classified by the gateway. Web search backs any team that needs current facts; a sandboxed Python tool runs untrusted code in a restricted subprocess; a read-only SQL tool is SELECT-only, allowlisted and row-limited against a replica.

Engineering · web · codeResearch & Analytics · web · code · SQLContent & Marketing · webCustomer & Sales · webOperations, Legal & Compliance · web · SQLQA gate

Deploy

Ship it with Docker Compose (Postgres + app, migrations applied on start), run it locally with an in-memory checkpointer and auth off, or drive it from the shinrai CLI. A Postgres-backed LangGraph checkpointer makes every conversation a durable, resumable thread.

Python 3.11LangGraphFastAPI + SSEPostgreSQLOAuth2 + JWTDocker

Missions & proposals

Two ways in. One governed workforce.

Run a mission directly on a thread, or go through the governed proposal lifecycle — scope a raw request into a costed proposal, accept it, deliver it, and revise it with feedback. Both paths run the same instrumented graph.

The governed path

The proposal lifecycle

A raw client request is scoped at intake into a costed proposal an operator can accept or reject. Accepting and delivering dispatches the full workforce on a dedicated, tracked thread. A delivered proposal can be revised with feedback, which dispatches a fresh run honouring the original scope plus the change.

A · IntakeB · DeliveryC · Revision
proposal statusLifecycle
intake bounded scoping pass
Scopedproposednot_readyclarifying
Decisionacceptedrejected
Deliverydeliveredrun linked
Revisionfeedbackfresh run

scoping and delivery run the workforce graph accept / reject / deliver by an operator

Pipeline A — Intake

  • A deliberately cheap, bounded scoping pass: readiness gate → plan → deliverables, metered for an estimated cost
  • A ready request becomes a proposed proposal with objective, deliverables and plan
  • A not-ready one becomes clarifying with an auto-generated clarification interview

Pipeline B — Delivery

  • Accepting and delivering dispatches the full workforce run through the run manager
  • Runs on a dedicated thread; the proposal moves to delivered and links its delivery run
  • Tracked as a cancellable task with its own per-run budget

Pipeline C — Revision

  • A delivered proposal can be revised with client feedback
  • Dispatches a fresh delivery run honouring the original scope plus the feedback
  • Every revision is linked back to the proposal it came from

Direct missions

  • Skip the proposal and run a mission straight on a thread
  • Autonomous end to end: plan → compose → execute → QA gate → finalize
  • Stream activity over SSE, or get the final answer in one call
Bounded

Hard LLM-call and wall-clock budgets per mission guarantee every run terminates.

Resumable

A Postgres checkpointer makes every thread durable — cancel, resume or restart a run.

Costed

Token accounting and an estimated USD cost from a governed price table, per run.

Multi-tenant

Threads are tenant-owned; checkpointer thread ids are namespaced per tenant.

The clarification interview

When scoping finds gaps, the platform raises a round of role-addressed questions — business, product owner, CTO, CISO, data protection, legal, finance, operations, end user — instead of guessing. Answers can be captured in the admin UI or exported as a Markdown, CSV or print-ready HTML questionnaire and re-imported. Every answer is attributed to the authenticated responder, and intake-origin rounds fold the answers back in and re-scope, bounded by a maximum number of rounds.

Role-addressedExport / importAttributable answers

The teams on the roster

A mission is staffed from a declarative org chart. Adding a team or an expert is pure data — the orchestration engine needs no changes.

web · code

Engineering

Full SDLC: product, systems analysis, stack strategy, UI/UX, architecture, backend, frontend, data engineering, SDET, QA, security, DevOps and docs.

web · code · SQL

Research & Analytics

Head of research, research analyst, data scientist, BI analyst and an insights synthesiser.

web

Content & Marketing

Content strategist, copywriter, editor and SEO specialist.

web

Customer & Sales

Support engineer, customer success manager and enterprise account executive.

web · SQL

Operations, Legal & Compliance

Compliance reviewer, operations analyst and the QA Reviewer that gates acceptance.

declarative

Your own teams

Declare a new team or expert — its remit, tool grants and artifact types — and the composer can staff it without touching the graph.

How Butai compares

A workforce, not a bigger prompt.

Assistants and frameworks help one model do more. Butai is a governed organisation of agents that plans, divides work, gates its own quality, and records how it got there.

QuestionSingle-agent assistantsAgent frameworksDIY orchestrationTypical AI consultancyBUTAI By SHINRAI TECHNOLOGIES LIMITED
Plans the mission & a definition of done◐Ad hoc◐You build it◐You build it◐In a document●Built in: objective + checkable definition of done
Hierarchical teams of specialists○One generalist◐Primitives only◐Hand-wired○Human-only●Chief → teams → world-expert members
Checks itself before finishing○No◐If you add it◐If you add it◐Varies●A QA gate others cannot skip
Governed models & tools○Raw model id○Your responsibility○Your responsibility◐Policy on paper●Model registry + risk-classified tool gateway
Records how the result was reached○Chat log◐Traces if wired◐If you build it◐Written afterwards●Queryable decision trace with unit economics
Guaranteed to terminate◐Context limit○Can loop○Can loop●Human-paced●Hard LLM-call & wall-clock budgets

A general comparison of approaches, not specific products. Capabilities of any tool vary and change over time.

Operate it

Dispatched, monitored, and under your control.

  1. 01 · Dispatch

    Every mission is a tracked run

    The run manager compiles a per-run graph with its own budget, trace recorder and cost accumulator, and tracks it as a cancellable task that heartbeats to the runs table.

  2. ◆ operator console02 · Monitor

    Down to the individual agent

    Watch each run’s status, current node and team · agent, step and call counts, budget usage, and the full artifact and handoff trail. A stale heartbeat flags a run as stuck.

  3. ◆ human decision03 · Control

    Cancel, resume, restart, clarify

    Resume continues the same thread from its last checkpoint, restart re-runs from the start, and clarify answers a run paused mid-mission and resumes it with the answer injected.

Who stands behind it

SHINRAI is the firm behind the platform.

Butai is SHINRAI’s own agentic workforce platform, and SHINRAI stands behind it. We are an AWS Advanced Tier Services Partner with offices in Nairobi and Dubai, delivering cloud, data and AI work for clients across industries.

AWS Advanced Tier Services PartnerSnowflake Select PartnerInformatica IDMC Partner
Founder & CEO

Timothy Munyao

Holds 13 AWS certifications and the AWS Golden Jacket, the first Kenyan to receive it. Leads Butai’s design and governance.

30%lower costs, about KES 1M saved a yearUnity Homes
40%fewer manual IT tasks, 20 hours saved a weekAdrian Group
83%faster deploymentsPallax Kenya
99.99%uptimePM Realty

Results from SHINRAI client engagements, as published on shinraitechnologies.io.

Fit

Who Butai is for.

A good fit

  • Teams who want multi-step work delivered by a crew of specialist agents, not a single chat
  • Organisations that need to see how an AI result was reached, and what it cost
  • Builders who want autonomy with a QA gate, budgets and governance already wired in
  • Multi-tenant deployments that need per-tenant isolation and auth from day one

Not a fit

  • A quick one-off answer where a single assistant is already enough
  • Fully unattended autonomy with no QA gate, budgets or human-in-the-loop
  • Work with no checkable definition of done anyone can agree on
  • Use cases that need a raw model id with no governance in the path

Questions

What teams ask first.

What is Butai, in one line?

An autonomous, hierarchical multi-agent platform. You give it a mission; it plans the objective and a definition of done, composes a crew of world-expert agents, executes through typed handoffs, gates its own quality, and finalises an answer — all recorded in a decision trace and bounded by budgets.

Is it fully autonomous?

It is autonomous within guardrails. An autonomy ladder (L0–L3) clamps how much it may do on its own, a QA gate decides when work is done, and budgets guarantee termination. When a mission is genuinely ambiguous, a run can pause mid-flight and ask a human, then resume with the answer injected.

How do the agents actually collaborate?

Through structure, not a growing chat. A Chief Supervisor routes among teams; each Team Supervisor routes to its experts. Experts publish typed, versioned artifacts to a shared blackboard and hand off explicitly, so each recipient gets clean, deterministic context.

Which models does it use?

Whichever your registry approves. Agents never reference a raw model id — they resolve a logical alias through a governed registry that enforces approval status, data tier and guardrails. Backends include a local manifest, AWS SageMaker and Snowflake Cortex, all sharing one enforcement path.

What stops it from running away or getting expensive?

Every mission runs under a hard cap on LLM calls and a wall-clock budget, checked at every hop. Each run gets its own budget so one stuck mission can’t drain the platform, and an optional hard deadline marks a run timed-out while leaving a checkpoint to resume from.

What’s the difference between a mission and a proposal?

A mission runs the workforce directly on a thread. A proposal is the governed path: a raw request is scoped at intake into a costed proposal an operator accepts or rejects, delivers as a tracked run, and can revise with feedback. Both run the same graph.

Can a human step in during a run?

Yes. On a genuine ambiguity the run raises a real interrupt that checkpoints the graph and holds as awaiting-clarification — not active, not terminal. An operator answers from the admin portal and the run resumes exactly where it paused, surviving a restart if needed.

Can we see how a result was produced?

Yes. Every mission writes a queryable decision trace — routing and rationale, tool calls, artifacts, QA verdicts, plan deviations and clarifications — plus token and cost totals. It streams live over SSE and is browsable in the admin portal down to the individual agent.

Can we add our own teams or tools?

Yes. The org chart is declarative data: a team declares its members, their tool grants and the artifact types they produce. Adding a team or expert needs no change to the orchestration engine, and new tools are registered in the gateway, which fails closed on anything unregistered.

Next step

Request a demo.

Tell us about the work you want a workforce to run. We’ll reply to arrange a 30-minute walkthrough on a mission of your choosing.

  1. Walkthrough. We run a mission end to end — plan, crew, execution, QA gate, trace.
  2. Your use case. We map your work to teams, tools and the registry backend that fits.
  3. Pilot. A scoped pilot on your own missions, with the decision trace to show for it.