MUTOWA · Model Documentation · Current release: Super Heavy 2 · October 2026
KOTO is our flagship frontier ensemble — a distributed network of 42 specialized 32B nodes coordinating in real time, powered entirely by solar energy across Canada, the United States, and Japan. The current lineup — KOTO Flash 1.2, KOTO Super 1.2, and KOTO Super Heavy 2 — runs on an ensemble that aggregates 1.34T parameters across its dedicated edge nodes for generating images, writing code, and making files.
Download this documentation (Markdown)
KOTO is MUTOWA's flagship frontier ensemble. Rather than a single monolithic network, KOTO distributes work across 42 specialized 32B nodes that coordinate in real time. Each node is sharpened for a domain — reasoning, code, vision, language, file construction — and an orchestration layer routes each request across the nodes best suited to it, so the ensemble behaves as one coherent model.
The result is a system that reasons deeply, responds fast, and produces finished artifacts: images, codebases, and files — end to end, from a single prompt.
| Metric | Value |
|---|---|
| Aggregate parameters | 1.34T across edge nodes |
| Context window | Up to 2,000,000 tokens (KOTO Super Heavy 2) · 1,000,000 (KOTO Super) · 128K (Flash) |
| Throughput | ~200 tokens per second (typical sustained, ensemble-wide) |
| Specialized nodes | 42 |
| Availability target | 99.999% |
| Power source | 100% solar |
KOTO runs on MUTOWA's Private Edge Compute network: solar-powered edge sites in Canada, the United States, and Japan, linked by millisecond anycast routing and post-quantum encrypted tunnels. Inference happens close to where requests originate — which is why KOTO is fast, private, and sustainable at the same time.
The ensemble is composed of 42 nodes, each a 32B-parameter specialist:
A request enters the ensemble once and is decomposed by a coordinator node into subtasks. Specialist nodes work in parallel where tasks are independent and in sequence where they depend on each other. Intermediate results pass between nodes as structured context, so the final answer is assembled — not averaged. This is a mixture-of-experts approach taken to the infrastructure level: expertise is a property of hardware placement, not just of model weights.
On the thinking tiers, KOTO plans before it answers: it decomposes the problem, works through it step by step, and verifies its own conclusion against the question before responding. The thinking budget scales with the tier you choose.
KOTO Super Heavy 2 accepts up to 2,000,000 tokens of context (1,000,000 on Super) — whole codebases, full document corpora, or hour-long transcripts — with retrieval-augmented grounding across the window. Over-long conversations are compacted automatically rather than rejected: the oldest turns are summarized away while the latest request is always preserved intact.
For ambiguous, safety-critical, or multiple-choice questions, KOTO does not bet on a single generation. Following the self-consistency technique (Wang et al., 2022), it produces one deterministic answer plus several independent sampled attempts, and the majority answer wins. If no answer reaches consensus, the ensemble escalates automatically — more samples, a larger thinking budget, stricter answer extraction — and the caller simply receives a better-reasoned response. The mechanics are never exposed to the client.
KOTO is served through three tiers on the public API. All tiers share the same ensemble, weights, and knowledge — they differ in how much thinking and node coordination a request mobilizes.
| Model id | Profile | Thinking | Context window | Output per run | Best for |
|---|---|---|---|---|---|
KOTO-flash-1.2 (alias KOTO-lite-1.2) | Fast, direct | Minimal | 128K tokens | 8K tokens | Everyday conversation, classification, extraction, high-throughput workloads |
KOTO-super-1.2 | Balanced, flagship | Extended | 1M tokens | 1M tokens | General-purpose assistant work: writing, analysis, multi-step tasks, code |
KOTO-super-heavy-2 | Deepest, maximum envelope | Always on, maximum budget | 2M tokens | 2M tokens | Hard reasoning, agentic planning, whole-repo comprehension, system architecture, research-grade questions |
Choosing a tier. Start with KOTO-super-1.2. Move down to KOTO-flash-1.2 when latency or volume dominates; move up to KOTO-super-heavy-2 when a task resists a first try — mathematics, long-horizon planning, subtle concurrency bugs — or when the input alone is a repository, a data dump, or a book.
Super Heavy 2 — the widest envelope in production. KOTO Super Heavy 2 carries a 2,000,000-token context window and a matching 2,000,000-token output budget per run — the only tier where what comes out can be as large as what goes in. Read an entire codebase, a legal discovery dump, or a book-length corpus in one prompt. Write a full novel, a complete platform, or an end-to-end build in one run. KOTO Super carries half: 1M in, 1M out. A single engine call is still bounded by the engine maximum; longer work is extended automatically across continuations until the run budget is reached, so the ceiling you can use is the one advertised. The legacy id
KOTO-super-heavy-1.0still resolves to the same tier.
Unknown or omitted model ids ride the flash tier — the API never errors on a model name.
Prompt to pixels — any style, any scale. KOTO's vision nodes generate finished artwork end-to-end: no retouching, no external editor. Diffusion pipelines run directly on the ensemble's GPU accelerators.
KOTO plans the architecture, writes the implementation, runs the tests, and fixes its own mistakes. One prompt in — a production-ready change out:
$ koto run "dashboard for womens world cup"
✓ Plan created — 4 tasks
✓ WebSocket server — Node.js, 42 lines
✓ Live metric cards — React 19
✓ Tests passed — 18/18
● deployed to preview in 38s
KOTO produces complete, verbatim multi-file artifacts — documents, data formats, structured exports — not sketches. Files are emitted with deterministic names and full contents.
KOTO can invoke tools (web search, retrieval) as part of answering, grounding real-time facts into its responses.
| Modality | Input | Output |
|---|---|---|
| Text | Yes | Yes |
| Code | Yes | Yes |
| Images | Yes | Yes (diffusion nodes) |
| Audio | Yes | — |
| Files | Yes | Yes |
Every KOTO node is a purpose-built edge computer, designed for silent, unattended, solar-powered operation.
| Component | Specification |
|---|---|
| Compute nodes | 42 |
| SoC | AMD EPYC Embedded 3251 — 8 cores / 16 threads, 2.5 GHz base, 3.1 GHz boost |
| GPU accelerators | Integrated edge-GPU module + 1× high-end discrete accelerator |
| System RAM | 128 GB (4×32 GB) DDR5-5600 ECC unbuffered SO-DIMM per node |
| Ephemeral memory | 32 GB RAMDisk — tmpfs mounted noexec, nosuid, nodev |
| Boot drive | 1 TB Samsung 990 Pro PCIe Gen4 NVMe (OS image and kernel) |
| Thermal design | Passively cooled, fanless aluminum chassis, 15 W–55 W configurable TDP |
| Isolation | AMD SEV-SNP encrypted virtualization with PSP-managed memory keys |
| Metric | Value |
|---|---|
| Total bandwidth | 16.8 Tbps |
| Average latency | < 1.2 ms |
| Node latency | P99 round trip under 1.2 ms across regional fiber links |
| Network peering | Tier-1 |
| Uptime | 99.999% |
Packets are classified and steered at the driver level before the kernel wakes up — the ensemble's coordination traffic never contends with the OS network stack.
Transpacific and transcontinental traffic rides dedicated fiber, including the FASTER and JUPITER submarine cable systems, with landings at Chikura and Shima (Japan) terminating to San Francisco and Los Angeles. Regional anycast routing keeps a request on the lowest-latency path between Canada, the United States, and Japan.
| Platform | Value |
|---|---|
| Protocols | gRPC, WebSocket |
| Official SDKs | TypeScript, Go, Rust |
| Public API wire formats | Standard chat-completions and messages shapes |
KOTO runs on 100% solar power. The ensemble's sovereign computational fabric is deployed across an isolated, dedicated 42-node bare-metal cluster operating on a private, microgrid-supported solar array.
| System attribute | Specification | Operational metric |
|---|---|---|
| Total compute nodes | 42 high-density multi-core heterogeneous nodes | 336 high-performance cores / 5.3 TB ECC RAM |
| Primary power source | Photovoltaic bifacial monocrystalline solar array | 145 kWp peak DC generating capacity |
| Energy storage | LiFePO₄ (lithium iron phosphate) microgrid BESS | 720 kWh dedicated reserve (~18 h autonomy) |
| Cluster PUE | Power Usage Effectiveness | 1.04 (direct negative-pressure ambient air heat exchange) |
| Lifecycle carbon | Scope 1 & 2 operational emissions | 0.00 g CO₂e / MegaToken (100% off-grid solar net-zero compute) |
| Cooling | Direct convective air, negative-pressure flow | Sub-micron HEPA filtration, vertical thermal-stack exhaust |
| Water consumption | — | Zero gallons of potable water during inference cycles |
Hyperscale data centers frequently mask their true environmental footprint behind carbon-offset certificates while consuming clean regional drinking water for evaporative cooling. MUTOWA's array rejects this approach: ambient outside air is drawn through sub-micron HEPA filtration arrays, circulated directly over the 42 compute chassis, and vented vertically via thermal stack effects. The cluster consumes exactly zero gallons of potable water during inference cycles, and the realized facility PUE of 1.04 compares against the enterprise data center average of 1.58.
The 42 nodes are managed by MUTOWA's proprietary Solar-Aware Load Director (SALD). Inference tasks are classified into two latency vectors:
The sustainability thesis is simple: clean power where the work happens. Millisecond compute does not require a mega-datacenter; it requires well-placed, efficient machines on renewable generation.
KOTO's security model is post-quantum by default and attested in hardware.
| Layer | Mechanism |
|---|---|
| Encryption in transit | NIST ML-KEM (Kyber-1024) key encapsulation, AES-256-GCM WireGuard tunnels |
| Perfect forward secrecy | Session keys rotate every 60 seconds via CPU true RNG |
| Encryption in use | AMD SEV-SNP encrypted virtualization, PSP-managed memory keys |
| Ephemeral state | 32 GB RAMDisk, mounted noexec / nosuid / nodev |
| Telemetry | Zero — requests are not logged, profiled, or used for training |
All inter-node and client-facing tunnels use Kyber-1024 key encapsulation layered under AES-256-GCM inside WireGuard. The mesh is therefore resistant to harvest-now-decrypt-later collection: traffic captured today remains unreadable to quantum adversaries later.
Every tunnel's session key is re-derived from the CPU hardware random number generator every 60 seconds. A compromised key exposes at most one minute of a single session.
Node workloads execute inside AMD SEV-SNP encrypted virtual machines with measured boot. Memory is encrypted with PSP-managed keys that never leave the silicon; remote attestation proves the exact code running before a node joins the ensemble.
MUTOWA operates Private Edge Compute with zero telemetry: KOTO does not retain your prompts, outputs, or usage fingerprints for its own purposes. Inference state lives in the noexec RAMDisk and evaporates on session end.
| Region | Nodes |
|---|---|
| Canada | 14 |
| United States | 20 |
| Japan | 8 |
| Total | 42 |
| Coverage | Value |
|---|---|
| Live cities | 7 |
| Countries | 3 |
The footprint follows MUTOWA's founding geography — Winnipeg, Canada and Tokyo, Japan — with dense United States coverage for transpacific peering. Anycast routing directs each request to the healthiest, nearest capacity; regional failures drain to neighboring sites automatically, which is where the 99.999% uptime figure comes from.
The KOTO API is served at https://api.mutowa.com. Endpoints are drop-in compatible with the two most common chat-API wire shapes, so most existing tooling connects by pointing it at the KOTO base URL with a KOTO key.
curl https://api.mutowa.com/v1/models
curl https://api.mutowa.com/v1/chat/completions \
-H "Authorization: Bearer mk_…" \
-H "content-type: application/json" \
-d '{
"model": "KOTO-super-1.2",
"messages": [
{"role": "system", "content": "You are a precise research assistant."},
{"role": "user", "content": "Summarize the state of fault-tolerant quantum error correction."}
]
}'
curl https://api.mutowa.com/v1/messages \
-H "x-api-key: mk_…" \
-H "content-type: application/json" \
-d '{
"model": "KOTO-super-heavy-1.0",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Design a rate limiter for a multi-region API."}]
}'
Set "stream": true on either wire format. Deltas arrive as server-sent events; streams are unbuffered end to end, so first-token latency reflects ensemble dispatch rather than proxy buffering.
All completion endpoints require a Bearer key — Authorization: Bearer mk_… (the messages shape also accepts x-api-key). Keys are issued by MUTOWA and can carry per-key rate limits (requests per minute, tokens per minute) and output caps.
| Endpoint | Shape | Description |
|---|---|---|
GET /v1/models | — | List the KOTO tiers |
GET /api/health | — | Service health and dependency status |
GET /docs | — | This documentation (HTML) |
GET /docs/download | — | This documentation (Markdown download) |
POST /v1/chat/completions | Chat | Chat completions; supports stream, temperature, top_p, max_tokens |
POST /v1/completions | Chat | Legacy text completions |
POST /v1/messages | Messages | Messages API; supports stream, max_tokens, system |
SDK base-URL alias of POST /v1/messages | Messages | Alias route for SDKs that pin a fixed base URL (exact spelling in the GET /v1 index) |
| Parameter | Type | Default | Description |
|---|---|---|---|
model | string | KOTO-flash-1.2 | Tier id. Unknown or omitted ids ride the flash tier — never an error. The retired KOTO-lite-1.2 and -contributor spellings are accepted as aliases. |
messages | array | required | Conversation turns. Over-long conversations are compacted automatically rather than rejected. |
temperature | number | 0 | Sampling temperature. The API defaults to deterministic output for reliability. |
top_p | number | — | Nucleus sampling. |
max_tokens | integer | platform default | Output cap per run: tier ceilings lite 8K · super 1M · super-heavy 2M (per-key caps only lower). Thinking headroom is added on top automatically on thinking tiers; a single engine call is clamped to the engine maximum and longer answers extend across automatic continuations. |
stream | boolean | false | Stream deltas over SSE. |
Errors return structured JSON — never internal traces.
| Status | Meaning |
|---|---|
400 | Malformed request (missing messages, invalid parameter) |
401 | Invalid or missing API key |
429 | Per-key rate limit or token budget exceeded (JSON body gives details) |
502 | Service momentarily saturated — safe to retry |
Background jobs and agent loops must never run away. KOTO enforces hard ceilings across all client and server boundaries; exceeding one surfaces as a structured error or a clean truncation — never a hang or an internal trace. The values below are platform defaults for this release and are enforced server-side:
| Subsystem | Hard limit boundary | Enforcement mechanism |
|---|---|---|
| Token burst ingress | 20,000 tokens per minute | Token-bucket rate limiter; queues requests exceeding ceiling |
| Active context buffer | Strictly 12 previous dialogue turns | Windowed truncation; prevents quadratic prompt inflation |
| Agent loop execution | Maximum 6 tool rounds per turn | Halts execution chain and presents collected findings |
| Web search operations | Maximum 2 rounds per query turn | Prevents infinite search loops on ambiguous queries |
| Subprocess code sandbox | 30-second timeout / 5 MB output limit | Immediate SIGKILL termination on ceiling breach |
| Local file I/O | ≤ 1 MB per atomic file operation | SafePath validator halts read/write on oversized payloads |
| Browser extraction | ≤ 100 KB raw DOM text (32 KB default) | Triggers clean Markdown extraction and truncates |
| ZIP compilation | ≤ 50 files / ≤ 5 MB total payload | In-memory buffer enforcement prior to disk stream |
| Mesh complexity | ≤ 20,000 total triangles | Procedural geometry-index ceiling to prevent GPU crash |
KOTO identifies as KOTO, an AI assistant made by MUTOWA. Its behavior contract:
About MUTOWA — from 無永遠 (mu eien), "boundless permanence" — MUTOWA is a technology company building enduring, fault-tolerant infrastructure across quantum-ready systems, applied AI (including KOTO), and private edge computing. Founded by Parker Czuba (Co-Founder & Chief Technology Director) and Nicolai Kazuyuki Diaz Tsuzuki (Co-Founder & President), operating from Winnipeg, Canada and Tokyo, Japan. Website: mutowa.com.
KOTO is trained with reinforcement learning from human feedback and a constitutional training regime, evaluated continuously for factuality, instruction fidelity, and refusal calibration. Alignment research runs in the loop across releases rather than being applied after the fact.
Toward the mission: every KOTO release narrows the gap between narrow tools and general intelligence — deeper reasoning, longer memory, richer action — safely, and on purpose.
| Area | Current lineup — Flash 1.2 · Super 1.2 · Super Heavy 2 |
|---|---|
| Ensemble | 42-node real-time coordination; 1.34T aggregate parameters across dedicated edge nodes |
| Reasoning | Extended chain-of-thought with tier-dependent thinking budgets |
| Vision | Dedicated diffusion edge nodes for end-to-end image generation |
| Agentic coding | Plan → implement → test → repair loop producing production-ready changes |
| Context | Tiered context windows — 2M tokens on Super Heavy 2, 1M on Super — with automatic conversation compaction |
| Output | Per-run output budgets to match the window: 2M tokens on Super Heavy 2, 1M on Super, delivered across automatic continuations |
| API | Three public tiers — KOTO-flash-1.2, KOTO-super-1.2, KOTO-super-heavy-2 (legacy KOTO-lite-1.2 and KOTO-super-heavy-1.0 aliased) — on standard chat-completions and messages endpoints |
| Infrastructure | Fully solar-powered footprint; post-quantum mesh (Kyber-1024) with 60-second key rotation; 99.999% uptime |
MUTOWA — from the Japanese 無永遠, boundless permanence — is a technology company linking quantum computing, applied artificial intelligence, and private edge computing: building foundational frameworks tailored for the next generation of computation.
| Name | Role |
|---|---|
| Nicolai Kazuyuki Diaz Tsuzuki | Co-Founder & President (代表取締役社長) |
| Parker Czuba | Co-Founder & Chief Technology Director (取締役最高技術責任者) |
Winnipeg, Canada · Tokyo, Japan — mutowa.com
The philosophy: precision, longevity, and discipline — infrastructure designed to endure.
KOTO · Technical Documentation · © 2026 MUTOWA. All specifications subject to change as the network evolves.