A Chief routes to teams; each team routes to experts.
A Chief Supervisor picks the next team. Each Team Supervisor picks the next world-expert member. Routing is a structured choice among the allowed options, not a free-for-all.
Autonomous · Hierarchical · Governed · Multi-tenant
Butai is SHINRAI’s agentic AI workforce. It plans the mission, composes a crew of world-expert agents, executes through typed handoffs, gates its own quality, and records every decision in a ledger you can read — all inside budgets that guarantee it stops.
| seq | event | actor | run |
|---|---|---|---|
| 01 | planobjective + definition of done | node:plan | 7a1c…e04 |
| 02 | composeengineering + research + QA | node:composer | 7a1c…e04 |
| 03 | route→ engineering | chief:supervisor | 7a1c…e04 |
| 04 | artifactspec published | eng:architect | 7a1c…e04 |
| 05 | tool_callcode · sandboxed | eng:backend | 7a1c…e04 |
| 06 | clarify◆ paused for a human answer | human:operator | 7a1c…e04 |
| 07 | handoffeng → research | chief:supervisor | 7a1c…e04 |
| 08 | artifactbenchmark analysis | res:data-scientist | 7a1c…e04 |
| 09 | qa_gate◆ verified vs definition of done | ops:qa-reviewer | 7a1c…e04 |
| 10 | finalizeanswer synthesised | chief:supervisor | 7a1c…e04 |
Why Butai exists
A single prompt can draft something plausible. Delivering a real objective takes a team that plans, divides the work, checks itself, and leaves a record of how it got there.
How a mission runs
Every mission follows the same path through a hierarchical LangGraph. Teams of world-expert agents do the work; a QA gate decides when it is finished; and the whole run is bounded so it always terminates.
Sharpens the mission into an objective and a checkable definition of done, and lists the teams and steps it will take.
Assembles the crew for this mission — from the plan or by relevance — and always includes the QA gate team.
Routes the mission among the composed teams, hop by hop, until the work is ready for review.
Each team supervisor routes to its world-expert members, who use their granted tools and publish typed artifacts to the shared blackboard.
If the mission is genuinely ambiguous, the run can pause mid-flight on a real interrupt, ask a person, and resume with the answer injected.
Experts hand off explicitly — within a team or across teams — carrying the artifacts the next recipient needs. No shared context by accident.
The Operations QA Reviewer verifies the work against the definition of done. Only an approved result finishes; a revision loops back, bounded.
The Chief synthesises the polished final answer from the approved artifacts and the mission closes.
Governance
Six enforcement seams sit under every mission. Each one is a boundary an agent cannot talk its way past. Choose a seam to see what it enforces and how.
No agent ever references a raw provider model id. Every LLM resolves a logical alias through a governed registry that enforces policy on the way through.
| What it enforces | How |
|---|---|
| Only approved models are used | Aliases like expert-default resolve to a model with an approval status; unapproved models are refused at resolution |
| Data-privacy tier is respected | Each alias carries a data-tier and guardrail metadata the enforcement path checks before returning a client |
| One governed path for every backend | Local manifest, AWS SageMaker Model Package Groups and Snowflake Cortex adapters share a single enforcement path — none can bypass it |
| Every resolution is auditable | Alias resolutions are logged, so you can see which governed model served which mission |
Every tool call passes through one gateway that classifies its risk and fails closed. Agents get capability by least privilege, not by default.
| What it enforces | How |
|---|---|
| Risk is explicit | Actions are classified READ < COMPUTE < NETWORK < WRITE < DESTRUCTIVE; a configurable ceiling blocks anything above it |
| Deny by default | An unregistered tool defaults to WRITE and is refused, so new capability must be registered deliberately |
| Least privilege per team | Web search, sandboxed code execution and read-only SQL are granted per team, not globally |
| Hard timeouts | Each tool call has a per-call hard timeout, so a hung tool cannot stall a run |
A single dial, L0 to L3, sets how much the workforce may do on its own. The effective policy is always the tighter of the level and the explicit settings.
| What it enforces | How |
|---|---|
| A clear ceiling per level | Each level clamps the tool-risk ceiling, whether the QA gate is required, and how many revisions are allowed |
| Tighter always wins | If the level and the explicit settings disagree, the stricter one applies — you cannot accidentally widen autonomy |
| Exceptions are recorded | Exceeding the ceiling requires an explicit, recorded waiver — never a silent escalation |
| QA gate on by default | At the default level, no mission finishes without passing the QA acceptance gate |
Autonomy without a stop condition is a liability. Every mission runs under hard and soft budgets that are checked at hop boundaries.
| What it enforces | How |
|---|---|
| A hard LLM-call cap | A per-mission maximum on total model calls that guarantees the run terminates |
| A wall-clock budget | A soft time budget checked at every hop, plus an optional hard deadline that marks a run timed-out and leaves a checkpoint to resume from |
| Per-run isolation | Each run gets its own budget, so one stuck mission cannot drain the shared budget of the platform |
| Stuck-run detection | Heartbeats to the runs table flag a mission as stuck when its heartbeat goes stale, for an operator to act on |
Butai is multi-tenant by construction. Conversations are isolated per tenant and access is gated by OAuth2 + JWT with MFA on privileged accounts.
| What it enforces | How |
|---|---|
| Tenant isolation | Tenant-scoped queries filter by tenant, and checkpointer thread ids are namespaced per tenant |
| Scoped access | JWTs carry scopes — chat for running missions, admin for the operator console |
| MFA on privileged accounts | TOTP is mandatory for admin-scoped accounts; a password alone mints only a short-lived challenge token |
| Fail-fast in production | Production startup refuses to run without a strong JWT secret and an explicit CORS allowlist |
Every mission writes a behavioural ledger — not a chat log — that an operator can query to see exactly how a result was reached.
| What it records | How |
|---|---|
| Routing & rationale | Each Chief and Team Supervisor decision, with the reason, attributed to a team and agent |
| Tool calls & artifacts | Every gateway-checked tool call and every typed artifact published to the blackboard |
| QA verdicts & clarifications | Acceptance decisions, revision loops, plan deviations and human clarification answers |
| Unit economics | Token accounting and an estimated USD cost from a governed price table, totalled per run |
These seams are enforced in code, not by policy documents. They are configurable per deployment and surfaced, with every admin action, in the operator console’s audit log.
Decision trace
$ shinrai plan "Design and benchmark a rate limiter, then review it." plan objective + definition of done, teams: engineering, research, operations $ shinrai compose "Research the top vector databases and recommend one." crew research · operations(qa) — qa gate always included
The trace is built from the run itself, not written from memory. It records every routing decision and its rationale, every gateway-checked tool call, every typed artifact published to the blackboard and the handoffs that carried them, the QA verdicts and revision loops, any plan deviations, and the human clarifications folded back in.
It also carries the unit economics: token accounting and an estimated cost from a governed price table, totalled per run. An operator reads all of it in the admin portal — down to the individual agent — or over the API.
| Surface | Where | What it shows |
|---|---|---|
| Live stream | SSE on the run endpoint | plan, chief routing, member activity, artifacts, QA verdict and the final answer, as they happen |
| Admin portal | The /admin console | Every run’s status, current team · agent, budget usage, the full artifact + handoff trail, and metrics |
| Decision ledger | The trace store | A queryable record of routing, tool calls, verdicts, deviations and clarifications for audit |
A trace shows how a result was reached and what it cost. It is the difference between an agent you hope worked and a workforce you can inspect.
How the work holds together
Collaboration is explicit. Experts don’t share a growing transcript — they publish typed artifacts and hand off deliberately, so every recipient gets clean, deterministic context.
A Chief Supervisor picks the next team. Each Team Supervisor picks the next world-expert member. Routing is a structured choice among the allowed options, not a free-for-all.
Engineering spans the full SDLC — product, spec, architecture, backend, frontend, data, SDET, QA, security, DevOps, docs — alongside Research, Content, Customer and Operations teams.
Each expert declares the artifact types it produces — requirements, spec, design, code, test plan, review, analysis — and publishes versioned artifacts to a shared blackboard.
Handoffs name the sender, recipient, intent and the artifacts passed — intra-team, cross-team, escalation or return — so context moves on purpose, not by accident.
A dedicated Operations QA Reviewer verifies the mission against its definition of done. Only an approved result finishes; a needs-revision verdict routes back, bounded.
When a mission is genuinely ambiguous, the run interrupts and checkpoints, an operator answers, and it resumes with the answer injected — surviving a restart if needed.
Models & tools
Agents resolve a logical model alias through a governed registry, and reach for tools through one gateway. Swap the backend to fit your stack — the enforcement path, and the agents, don’t change.
| Registry backend | Resolves aliases to | Governance read from | Needs a cloud? |
|---|---|---|---|
| Local manifest | Models in a YAML/JSON manifest | Approval status, data tier and guardrails in the manifest | No — runs fully local |
| AWS SageMaker | Model Package Groups (latest Approved package) | Package approval status and governance metadata | AWS account |
| Snowflake Cortex | Models in the Snowflake Model Registry, served via Cortex | Version metadata on the registered model | Snowflake account |
Tools are granted per team by least privilege, and every call is risk-classified by the gateway. Web search backs any team that needs current facts; a sandboxed Python tool runs untrusted code in a restricted subprocess; a read-only SQL tool is SELECT-only, allowlisted and row-limited against a replica.
Ship it with Docker Compose (Postgres + app, migrations applied on start), run it locally with an in-memory checkpointer and auth off, or drive it from the shinrai CLI. A Postgres-backed LangGraph checkpointer makes every conversation a durable, resumable thread.
Missions & proposals
Run a mission directly on a thread, or go through the governed proposal lifecycle — scope a raw request into a costed proposal, accept it, deliver it, and revise it with feedback. Both paths run the same instrumented graph.
The governed path
A raw client request is scoped at intake into a costed proposal an operator can accept or reject. Accepting and delivering dispatches the full workforce on a dedicated, tracked thread. A delivered proposal can be revised with feedback, which dispatches a fresh run honouring the original scope plus the change.
scoping and delivery run the workforce graph accept / reject / deliver by an operator
proposed proposal with objective, deliverables and planclarifying with an auto-generated clarification interviewdelivered and links its delivery runHard LLM-call and wall-clock budgets per mission guarantee every run terminates.
A Postgres checkpointer makes every thread durable — cancel, resume or restart a run.
Token accounting and an estimated USD cost from a governed price table, per run.
Threads are tenant-owned; checkpointer thread ids are namespaced per tenant.
When scoping finds gaps, the platform raises a round of role-addressed questions — business, product owner, CTO, CISO, data protection, legal, finance, operations, end user — instead of guessing. Answers can be captured in the admin UI or exported as a Markdown, CSV or print-ready HTML questionnaire and re-imported. Every answer is attributed to the authenticated responder, and intake-origin rounds fold the answers back in and re-scope, bounded by a maximum number of rounds.
A mission is staffed from a declarative org chart. Adding a team or an expert is pure data — the orchestration engine needs no changes.
Full SDLC: product, systems analysis, stack strategy, UI/UX, architecture, backend, frontend, data engineering, SDET, QA, security, DevOps and docs.
Head of research, research analyst, data scientist, BI analyst and an insights synthesiser.
Content strategist, copywriter, editor and SEO specialist.
Support engineer, customer success manager and enterprise account executive.
Compliance reviewer, operations analyst and the QA Reviewer that gates acceptance.
Declare a new team or expert — its remit, tool grants and artifact types — and the composer can staff it without touching the graph.
How Butai compares
Assistants and frameworks help one model do more. Butai is a governed organisation of agents that plans, divides work, gates its own quality, and records how it got there.
| Question | Single-agent assistants | Agent frameworks | DIY orchestration | Typical AI consultancy | BUTAI By SHINRAI TECHNOLOGIES LIMITED |
|---|---|---|---|---|---|
| Plans the mission & a definition of done | ◐Ad hoc | ◐You build it | ◐You build it | ◐In a document | ●Built in: objective + checkable definition of done |
| Hierarchical teams of specialists | ○One generalist | ◐Primitives only | ◐Hand-wired | ○Human-only | ●Chief → teams → world-expert members |
| Checks itself before finishing | ○No | ◐If you add it | ◐If you add it | ◐Varies | ●A QA gate others cannot skip |
| Governed models & tools | ○Raw model id | ○Your responsibility | ○Your responsibility | ◐Policy on paper | ●Model registry + risk-classified tool gateway |
| Records how the result was reached | ○Chat log | ◐Traces if wired | ◐If you build it | ◐Written afterwards | ●Queryable decision trace with unit economics |
| Guaranteed to terminate | ◐Context limit | ○Can loop | ○Can loop | ●Human-paced | ●Hard LLM-call & wall-clock budgets |
A general comparison of approaches, not specific products. Capabilities of any tool vary and change over time.
Operate it
The run manager compiles a per-run graph with its own budget, trace recorder and cost accumulator, and tracks it as a cancellable task that heartbeats to the runs table.
Watch each run’s status, current node and team · agent, step and call counts, budget usage, and the full artifact and handoff trail. A stale heartbeat flags a run as stuck.
Resume continues the same thread from its last checkpoint, restart re-runs from the start, and clarify answers a run paused mid-mission and resumes it with the answer injected.
Who stands behind it
Butai is SHINRAI’s own agentic workforce platform, and SHINRAI stands behind it. We are an AWS Advanced Tier Services Partner with offices in Nairobi and Dubai, delivering cloud, data and AI work for clients across industries.
Holds 13 AWS certifications and the AWS Golden Jacket, the first Kenyan to receive it. Leads Butai’s design and governance.
Results from SHINRAI client engagements, as published on shinraitechnologies.io.
Fit
Questions
An autonomous, hierarchical multi-agent platform. You give it a mission; it plans the objective and a definition of done, composes a crew of world-expert agents, executes through typed handoffs, gates its own quality, and finalises an answer — all recorded in a decision trace and bounded by budgets.
It is autonomous within guardrails. An autonomy ladder (L0–L3) clamps how much it may do on its own, a QA gate decides when work is done, and budgets guarantee termination. When a mission is genuinely ambiguous, a run can pause mid-flight and ask a human, then resume with the answer injected.
Through structure, not a growing chat. A Chief Supervisor routes among teams; each Team Supervisor routes to its experts. Experts publish typed, versioned artifacts to a shared blackboard and hand off explicitly, so each recipient gets clean, deterministic context.
Whichever your registry approves. Agents never reference a raw model id — they resolve a logical alias through a governed registry that enforces approval status, data tier and guardrails. Backends include a local manifest, AWS SageMaker and Snowflake Cortex, all sharing one enforcement path.
Every mission runs under a hard cap on LLM calls and a wall-clock budget, checked at every hop. Each run gets its own budget so one stuck mission can’t drain the platform, and an optional hard deadline marks a run timed-out while leaving a checkpoint to resume from.
A mission runs the workforce directly on a thread. A proposal is the governed path: a raw request is scoped at intake into a costed proposal an operator accepts or rejects, delivers as a tracked run, and can revise with feedback. Both run the same graph.
Yes. On a genuine ambiguity the run raises a real interrupt that checkpoints the graph and holds as awaiting-clarification — not active, not terminal. An operator answers from the admin portal and the run resumes exactly where it paused, surviving a restart if needed.
Yes. Every mission writes a queryable decision trace — routing and rationale, tool calls, artifacts, QA verdicts, plan deviations and clarifications — plus token and cost totals. It streams live over SSE and is browsable in the admin portal down to the individual agent.
Yes. The org chart is declarative data: a team declares its members, their tool grants and the artifact types they produce. Adding a team or expert needs no change to the orchestration engine, and new tools are registered in the gateway, which fails closed on anything unregistered.
Next step
Tell us about the work you want a workforce to run. We’ll reply to arrange a 30-minute walkthrough on a mission of your choosing.