# KOTO — Technical Documentation

**MUTOWA · Model Documentation · Current release: Super Heavy 2 · October 2026**

KOTO is our flagship frontier ensemble — a distributed network of 42 specialized 32B nodes coordinating in real time, powered entirely by solar energy across Canada, the United States, and Japan. The current lineup — **KOTO Flash 1.2**, **KOTO Super 1.2**, and **KOTO Super Heavy 2** — runs on an ensemble that aggregates **1.34T parameters** across its dedicated edge nodes for generating images, writing code, and making files.

**Download & links**

- Browsable documentation: `https://api.mutowa.com/docs`
- This file: `https://api.mutowa.com/docs/download`
- API base URL: `https://api.mutowa.com`
- Company: [https://mutowa.com](https://mutowa.com)

---

## Contents

1. [Overview](#1-overview)
2. [The Ensemble — how KOTO works](#2-the-ensemble--how-koto-works)
3. [Model tiers](#3-model-tiers)
4. [Capabilities and modalities](#4-capabilities-and-modalities)
5. [Compute platform](#5-compute-platform)
6. [Network fabric](#6-network-fabric)
7. [Power and sustainability](#7-power-and-sustainability)
8. [Security architecture](#8-security-architecture)
9. [Global footprint](#9-global-footprint)
10. [API quick start](#10-api-quick-start)
11. [API reference](#11-api-reference)
12. [Model behavior and identity](#12-model-behavior-and-identity)
13. [Alignment and safety](#13-alignment-and-safety)
14. [Release notes](#14-release-notes)
15. [About MUTOWA](#15-about-mutowa)

---

## 1. Overview

KOTO is MUTOWA's flagship frontier ensemble. Rather than a single monolithic network, KOTO distributes work across **42 specialized 32B nodes** that coordinate in real time. Each node is sharpened for a domain — reasoning, code, vision, language, file construction — and an orchestration layer routes each request across the nodes best suited to it, so the ensemble behaves as one coherent model.

The result is a system that reasons deeply, responds fast, and produces finished artifacts: images, codebases, and files — end to end, from a single prompt.

**Headline numbers**

| Metric | Value |
|---|---|
| Aggregate parameters | 1.34T across edge nodes |
| Context window | 1,000,000 tokens |
| Throughput | ~200 tokens per second (typical sustained, ensemble-wide) |
| Specialized nodes | 42 |
| Availability target | 99.999% |
| Power source | 100% solar |

KOTO runs on MUTOWA's Private Edge Compute network: solar-powered edge sites in Canada, the United States, and Japan, linked by millisecond anycast routing and post-quantum encrypted tunnels. Inference happens close to where requests originate — which is why KOTO is fast, private, and sustainable at the same time.

## 2. The Ensemble — how KOTO works

### 2.1 Specialized nodes

The ensemble is composed of 42 nodes, each a 32B-parameter specialist. Node specializations include:

- **Reasoning nodes** — extended chain-of-thought planning, verification, and self-checking
- **Code nodes** — program synthesis, refactoring, test generation, and repair
- **Vision nodes** — image understanding and diffusion-based image generation
- **Language nodes** — long-form writing, translation, summarization
- **File nodes** — structured, multi-file artifact construction
- **Retrieval nodes** — in-context knowledge grounding

### 2.2 Real-time coordination

A request enters the ensemble once and is decomposed by a coordinator node into subtasks. Specialist nodes work in parallel where tasks are independent and in sequence where they depend on each other. Intermediate results pass between nodes as structured context, so the final answer is assembled — not averaged. This is a mixture-of-experts approach taken to the infrastructure level: expertise is a property of hardware placement, not just of model weights.

### 2.3 Extended chain-of-thought

On the thinking tiers, KOTO plans before it answers: it decomposes the problem, works through it step by step, and verifies its own conclusion against the question before responding. The thinking budget scales with the tier you choose (see [Model tiers](#3-model-tiers)).

### 2.4 Long context

KOTO Super Heavy 2 accepts up to **2,000,000 tokens** of context (1,000,000 on Super) — whole codebases, full document corpora, or hour-long transcripts — with retrieval-augmented grounding across the window. Over-long conversations are compacted automatically rather than rejected: the oldest turns are summarized away while the latest request is always preserved intact.

### 2.5 Self-consistency sampling

For ambiguous, safety-critical, or multiple-choice questions, KOTO does not bet on a single generation. Following the self-consistency technique (Wang et al., 2022), it produces one deterministic answer plus several independent sampled attempts, and the majority answer wins. If no answer reaches consensus, the ensemble escalates automatically — more samples, a larger thinking budget, stricter answer extraction — and the caller simply receives a better-reasoned response. The mechanics are never exposed to the client.

## 3. Model tiers

KOTO is served through three tiers on the public API. All tiers share the same ensemble, weights, and knowledge — they differ in how much thinking and node coordination a request mobilizes.

| Model id | Profile | Thinking | Context window | Output per run | Best for |
|---|---|---|---|---|---|
| `KOTO-flash-1.2` (alias `KOTO-lite-1.2`) | Fast, direct | Minimal | 128K tokens | 8K tokens | Everyday conversation, classification, extraction, high-throughput workloads |
| `KOTO-super-1.2` | Balanced, flagship | Extended | 1M tokens | 1M tokens | General-purpose assistant work: writing, analysis, multi-step tasks, code |
| `KOTO-super-heavy-2` | Deepest, maximum envelope | Always on, maximum budget | **2M tokens** | **2M tokens** | Hard reasoning, agentic planning, whole-repo comprehension, system architecture, research-grade questions |

> **Super Heavy 2 — the widest envelope in production.** KOTO Super Heavy 2
> carries a **2,000,000-token context window** and a matching **2,000,000-token
> output budget per run** — the only tier where what comes out can be as large
> as what goes in. Read an entire codebase, a legal discovery dump, or a
> book-length corpus in one prompt. Write a full novel, a complete platform,
> or an end-to-end build in one run. KOTO Super carries half: 1M in, 1M out. A
> single engine call is still bounded by the engine maximum; longer work is
> extended automatically across continuations until the run budget is reached,
> so the ceiling you can *use* is the one advertised. The legacy id
> `KOTO-super-heavy-1.0` still resolves to the same tier.

**Choosing a tier.** Start with `KOTO-super-1.2`. Move down to `KOTO-flash-1.2` when latency or volume dominates; move up to `KOTO-super-heavy-2` when a task resists a first try — mathematics, long-horizon planning, subtle concurrency bugs — or when the input alone is a repository, a data dump, or a book.

Unknown or omitted model ids ride the lite tier — the API never errors on a model name.

## 4. Capabilities and modalities

### 4.1 Image generation

Prompt to pixels — any style, any scale. KOTO's vision nodes generate finished artwork end-to-end: no retouching, no external editor. Diffusion pipelines run directly on the ensemble's GPU accelerators.

### 4.2 Agentic coding

KOTO plans the architecture, writes the implementation, runs the tests, and fixes its own mistakes. One prompt in — a production-ready change out. A typical session:

```text
$ koto run "dashboard for womens world cup"

✓ Plan created — 4 tasks
✓ WebSocket server — Node.js, 42 lines
✓ Live metric cards — React 19
✓ Tests passed — 18/18
● deployed to preview in 38s
```

### 4.3 File construction

KOTO produces complete, verbatim multi-file artifacts — documents, data formats, structured exports — not sketches. Files are emitted with deterministic names and full contents.

### 4.4 Tool use and retrieval

KOTO can invoke tools (web search, retrieval) as part of answering, grounding real-time facts into its responses.

### 4.5 Modality support

| Modality | Input | Output |
|---|---|---|
| Text | Yes | Yes |
| Code | Yes | Yes |
| Images | Yes | Yes (diffusion nodes) |
| Audio | Yes | — |
| Files | Yes | Yes |

## 5. Compute platform

Every KOTO node is a purpose-built edge computer, designed for silent, unattended, solar-powered operation.

### 5.1 Node hardware

| Component | Specification |
|---|---|
| Compute nodes | 42 |
| SoC | AMD EPYC Embedded 3251 — 8 cores / 16 threads, 2.5 GHz base, 3.1 GHz boost |
| GPU accelerators | Integrated edge-GPU module + 1× high-end discrete accelerator |
| System RAM | 128 GB (4×32 GB) DDR5-5600 ECC unbuffered SO-DIMM per node |
| Ephemeral memory | 32 GB RAMDisk — `tmpfs` mounted `noexec`, `nosuid`, `nodev` |
| Boot drive | 1 TB Samsung 990 Pro PCIe Gen4 NVMe (OS image and kernel) |
| Thermal design | Passively cooled, fanless aluminum chassis, 15 W–55 W configurable TDP |
| Isolation | AMD SEV-SNP encrypted virtualization with PSP-managed memory keys |

### 5.2 Why it looks like this

- **Fanless, wide-TDP passive cooling** means nodes live in sealed enclosures with no moving parts — a requirement for solar sites where maintenance visits are rare and dust, weather, and silence matter.
- **ECC memory** keeps month-long unattended uptimes free of silent corruption.
- **The noexec RAMDisk** absorbs all transient inference state; nothing written during a request survives a reboot, and nothing executable can be planted there.
- **SEV-SNP** means node workloads are encrypted in use — even a physical attacker with the chassis cannot read the memory of a running request.

## 6. Network fabric

### 6.1 Capacity and latency

| Metric | Value |
|---|---|
| Total bandwidth | 16.8 Tbps |
| Average latency | < 1.2 ms |
| Node latency | P99 round trip under 1.2 ms across regional fiber links |
| Network peering | Tier-1 |
| Uptime | 99.999% |

### 6.2 Per-node networking

- **NIC:** dual-port 400 GbE QSFP-DD PCIe 5.0 x16 controller with SR-IOV
- **Kernel bypass:** eBPF XDP direct driver-queue packet routing in under 40 ns

Packets are classified and steered at the driver level before the kernel wakes up — the ensemble's coordination traffic never contends with the OS network stack.

### 6.3 Long-haul links

Transpacific and transcontinental traffic rides dedicated fiber, including the **FASTER** and **JUPITER** submarine cable systems, with landings at Chikura and Shima (Japan) terminating to San Francisco and Los Angeles. Regional anycast routing keeps a request on the lowest-latency path between Canada, the United States, and Japan.

### 6.4 Protocols and SDKs

| | |
|---|---|
| Protocols | HTTPS REST (SSE streaming) — gRPC/WebSocket are roadmap, not shipped |
| Official SDKs | Any OpenAI-compatible or messages-shape client works as-is; first-party TS/Go/Rust SDKs are roadmap |
| Public API wire formats | Standard chat-completions and messages shapes |

## 7. Power and sustainability

**KOTO runs on 100% solar power.** The ensemble's sovereign computational fabric is deployed across an isolated, dedicated 42-node bare-metal cluster operating on a private, microgrid-supported solar array.

### 7.1 The sovereign solar cluster

| System attribute | Specification | Operational metric |
|---|---|---|
| Total compute nodes | 42 high-density multi-core heterogeneous nodes | 336 high-performance cores / 5.3 TB ECC RAM |
| Primary power source | Photovoltaic bifacial monocrystalline solar array | 145 kWp peak DC generating capacity |
| Energy storage | LiFePO₄ (lithium iron phosphate) microgrid BESS | 720 kWh dedicated reserve (~18 h autonomy) |
| Cluster PUE | Power Usage Effectiveness | **1.04** (direct negative-pressure ambient air heat exchange) |
| Lifecycle carbon | Scope 1 & 2 operational emissions | **0.00 g CO₂e / MegaToken** (100% off-grid solar net-zero compute) |
| Cooling | Direct convective air, negative-pressure flow | Sub-micron HEPA filtration, vertical thermal-stack exhaust |
| Water consumption | — | **Zero gallons** of potable water during inference cycles |

### 7.2 Why direct air cooling

Hyperscale data centers frequently mask their true environmental footprint behind carbon-offset certificates while consuming clean regional drinking water for evaporative cooling. MUTOWA's array rejects this approach: ambient outside air is drawn through sub-micron HEPA filtration arrays, circulated directly over the 42 compute chassis, and vented vertically via thermal stack effects. The cluster consumes exactly zero gallons of potable water during inference cycles, and the realized facility PUE of **1.04** compares against the enterprise data center average of 1.58.

### 7.3 Solar-aware compute pacing (SALD)

The 42 nodes are managed by MUTOWA's proprietary **Solar-Aware Load Director (SALD)**. Inference tasks are classified into two latency vectors:

- **Synchronous Interactive Stream (priority)** — real-time tokens for user chat and tool calls are served instantaneously from the combined output of the real-time PV array and the high-discharge LiFePO₄ battery reserve.
- **Deferred Asynchronous Syntheses (batch)** — offline workflows, background agent scheduled jobs, deterministic 3D mesh rendering, and parameter fine-tuning are scheduled dynamically against the real-time insolation curve. When solar irradiance exceeds **750 W/m²**, the cluster unlocks peak token throughput across all 42 nodes without grid interconnect draws.

### 7.4 Efficiency by design

- **Local generation, local compute.** Inference happens where the sun hits the panel — work is placed where clean power exists rather than moving data to distant datacenters. Less data moved is less energy spent.
- **Efficient hardware.** Fanless 15–55 W configurable-TDP nodes, passive aluminum cooling, and efficient accelerator silicon keep the draw inside what on-site generation provides.
- **N+2 UPS redundancy.** Battery banks ride through night hours and weather; two independent UPS units beyond requirement stand between any node and an interruption. Measured uptime: 99.999%.
- **The Sunlight Node program.** MUTOWA publishes the solar edge stack as licensable infrastructure — the Sunlight Node program (1,000-license single batch) lets operators run Private Edge Compute sites and earn from the compute they provide. The KOTO ensemble was the first large workload proven on this architecture.

The sustainability thesis is simple: **clean power where the work happens.** Millisecond compute does not require a mega-datacenter; it requires well-placed, efficient machines on renewable generation.

## 8. Security architecture

KOTO's security model is post-quantum by default and attested in hardware.

| Layer | Mechanism |
|---|---|
| Encryption in transit | NIST ML-KEM (Kyber-1024) key encapsulation, AES-256-GCM WireGuard tunnels |
| Perfect forward secrecy | Session keys rotate every 60 seconds via CPU true RNG |
| Encryption in use | AMD SEV-SNP encrypted virtualization, PSP-managed memory keys |
| Ephemeral state | 32 GB RAMDisk, mounted `noexec`/`nosuid`/`nodev` |
| Telemetry | Zero — requests are not logged, profiled, or used for training |

### 8.1 Post-quantum mesh

All inter-node and client-facing tunnels use Kyber-1024 key encapsulation layered under AES-256-GCM inside WireGuard. The mesh is therefore resistant to harvest-now-decrypt-later collection: traffic captured today remains unreadable to quantum adversaries later.

### 8.2 Sixty-second key rotation

Every tunnel's session key is re-derived from the CPU hardware random number generator every 60 seconds. A compromised key exposes at most one minute of a single session.

### 8.3 Attested enclaves

Node workloads execute inside AMD SEV-SNP encrypted virtual machines with measured boot. Memory is encrypted with PSP-managed keys that never leave the silicon; remote attestation proves the exact code running before a node joins the ensemble.

### 8.4 Privacy posture

MUTOWA operates Private Edge Compute with **zero telemetry**: KOTO does not retain your prompts, outputs, or usage fingerprints for its own purposes. Inference state lives in the noexec RAMDisk and evaporates on session end.

## 9. Global footprint

| Region | Nodes |
|---|---|
| Canada | 14 |
| United States | 20 |
| Japan | 8 |
| **Total** | **42** |

| Coverage | |
|---|---|
| Live cities | 7 |
| Countries | 3 |

The footprint follows MUTOWA's founding geography — Winnipeg, Canada and Tokyo, Japan — with dense United States coverage for transpacific peering. Anycast routing directs each request to the healthiest, nearest capacity; regional failures drain to neighboring sites automatically, which is where the 99.999% uptime figure comes from.

## 10. API quick start

The KOTO API is served at `https://api.mutowa.com`. Endpoints are drop-in compatible with the two most common chat-API wire shapes, so most existing tooling connects by pointing it at the KOTO base URL with a KOTO key.

### 10.1 List models

```bash
curl https://api.mutowa.com/v1/models
```

### 10.2 Chat-completions endpoint (curl)

```bash
curl https://api.mutowa.com/v1/chat/completions \
  -H "Authorization: Bearer mk_…" \
  -H "content-type: application/json" \
  -d '{
    "model": "KOTO-super-1.2",
    "messages": [
      {"role": "system", "content": "You are a precise research assistant."},
      {"role": "user", "content": "Summarize the state of fault-tolerant quantum error correction."}
    ]
  }'
```

### 10.3 Messages endpoint (curl)

```bash
curl https://api.mutowa.com/v1/messages \
  -H "x-api-key: mk_…" \
  -H "content-type: application/json" \
  -d '{
    "model": "KOTO-super-heavy-1.0",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "Design a rate limiter for a multi-region API."}]
  }'
```

### 10.4 Streaming (SSE)

Set `"stream": true` on either wire format. Deltas arrive as server-sent events; streams are unbuffered end to end, so first-token latency reflects ensemble dispatch rather than proxy buffering.

### 10.5 Authentication

All completion endpoints require a Bearer key — `Authorization: Bearer mk_…` (the messages shape also accepts `x-api-key`). Keys are issued by MUTOWA and can carry per-key rate limits (requests per minute, tokens per minute) and output caps.

## 11. API reference

### 11.1 Endpoints

| Endpoint | Shape | Description |
|---|---|---|
| `GET /v1/models` | — | List the KOTO tiers |
| `GET /api/health` | — | Service health and dependency status |
| `GET /docs` | — | This documentation (HTML) |
| `GET /docs/download` | — | This documentation (Markdown download) |
| `POST /v1/chat/completions` | Chat | Chat completions; supports `stream`, `temperature`, `top_p`, `max_tokens` |
| `POST /v1/completions` | Chat | Legacy text completions |
| `POST /v1/messages` | Messages | Messages API; supports `stream`, `max_tokens`, `system` |
| SDK base-URL alias of `POST /v1/messages` | Messages | Alias route for SDKs that pin a fixed base URL (exact spelling in the `GET /v1` index) |

### 11.2 Request parameters

| Parameter | Type | Default | Description |
|---|---|---|---|
| `model` | string | `KOTO-flash-1.2` | Tier id. Unknown or omitted ids ride the flash tier — never an error. The retired `KOTO-lite-1.2` and `-contributor` spellings are accepted as aliases. |
| `messages` | array | required | Conversation turns. Over-long conversations are compacted automatically rather than rejected. |
| `temperature` | number | `0` | Sampling temperature. The API defaults to deterministic output for reliability. |
| `top_p` | number | — | Nucleus sampling. |
| `max_tokens` | integer | platform default | Output cap per run: tier ceilings lite 8K · super 1M · super-heavy 2M (per-key caps only lower). Thinking headroom is added on top automatically on thinking tiers; a single engine call is clamped to the engine maximum and longer answers extend across automatic continuations. |
| `stream` | boolean | `false` | Stream deltas over SSE. |

### 11.3 Errors

Errors return structured JSON — never internal traces.

| Status | Meaning |
|---|---|
| `400` | Malformed request (missing `messages`, invalid parameter) |
| `401` | Invalid or missing API key |
| `429` | Per-key rate limit or token budget exceeded (JSON body gives details) |
| `502` | Service momentarily saturated — safe to retry |

### 11.4 Operational guardrails

Background jobs and agent loops must never run away. KOTO enforces hard ceilings across all client and server boundaries; exceeding one surfaces as a structured error or a clean truncation — never a hang or an internal trace. The values below are platform defaults for this release and are enforced server-side:

| Subsystem | Hard limit boundary | Enforcement mechanism |
|---|---|---|
| Token burst ingress | 20,000 tokens per minute | Token-bucket rate limiter; queues requests exceeding ceiling |
| Active context buffer | Strictly 12 previous dialogue turns | Windowed truncation; prevents quadratic prompt inflation |
| Agent loop execution | Maximum 6 tool rounds per turn | Halts execution chain and presents collected findings |
| Web search operations | Maximum 2 rounds per query turn | Prevents infinite search loops on ambiguous queries |
| Subprocess code sandbox | 30-second timeout / 5 MB output limit | Immediate SIGKILL termination on ceiling breach |
| Local file I/O | ≤ 1 MB per atomic file operation | SafePath validator halts read/write on oversized payloads |
| Browser extraction | ≤ 100 KB raw DOM text (32 KB default) | Triggers clean Markdown extraction and truncates |
| ZIP compilation | ≤ 50 files / ≤ 5 MB total payload | In-memory buffer enforcement prior to disk stream |
| Mesh complexity | ≤ 20,000 total triangles | Procedural geometry-index ceiling to prevent GPU crash |

## 12. Model behavior and identity

KOTO identifies as **KOTO, an AI assistant made by MUTOWA**. Its behavior contract:

- Answer the question that was actually asked — directly, completely, in the fewest words that stay clear.
- Lead with the answer; show reasoning only when the task needs it.
- Follow explicit formatting contracts (JSON, tables, word counts) exactly, even against its defaults.
- Grounds every claim: when it doesn't know, it says so plainly rather than filling gaps.

**About MUTOWA** — from 無永遠 (*mu eien*), "boundless permanence" — MUTOWA is a technology company building enduring, fault-tolerant infrastructure across quantum-ready systems, applied AI (including KOTO), and private edge computing. Founded by **Parker Czuba** (Co-Founder & Chief Technology Director) and **Nicolai Kazuyuki Diaz Tsuzuki** (Co-Founder & President), operating from Winnipeg, Canada and Tokyo, Japan. Website: [mutowa.com](https://mutowa.com).

## 13. Alignment and safety

KOTO is trained with reinforcement learning from human feedback and a constitutional training regime, evaluated continuously for factuality, instruction fidelity, and refusal calibration. Alignment research runs in the loop across releases rather than being applied after the fact.

Toward the mission: every KOTO release narrows the gap between narrow tools and general intelligence — deeper reasoning, longer memory, richer action — **safely, and on purpose**.

## 14. Release notes

| Area | Current lineup — Flash 1.2 · Super 1.2 · Super Heavy 2 |
|---|---|
| Ensemble | 42-node real-time coordination; 1.34T aggregate parameters across dedicated edge nodes |
| Reasoning | Extended chain-of-thought with tier-dependent thinking budgets |
| Vision | Dedicated diffusion edge nodes for end-to-end image generation |
| Agentic coding | Plan → implement → test → repair loop producing production-ready changes |
| Context | Tiered context windows — 2M tokens on Super Heavy 2, 1M on Super — with automatic conversation compaction |
| Output | Per-run output budgets to match the window: 2M tokens on Super Heavy 2, 1M on Super, delivered across automatic continuations |
| API | Three public tiers — `KOTO-flash-1.2`, `KOTO-super-1.2`, `KOTO-super-heavy-2` (legacy `KOTO-lite-1.2` and `KOTO-super-heavy-1.0` aliased) — on standard chat-completions and messages endpoints |
| Infrastructure | Fully solar-powered footprint; post-quantum mesh (Kyber-1024) with 60-second key rotation; 99.999% uptime |

### Roadmap

- **Sunlight Node** — the licensable solar edge node program expanding the network the ensemble runs on.

## 15. About MUTOWA

MUTOWA — from the Japanese 無永遠, *boundless permanence* — is a technology company linking quantum computing, applied artificial intelligence, and private edge computing: building foundational frameworks tailored for the next generation of computation.

**Three connected practices**

- **Private Edge Compute** — millisecond anycast routing across Canada, Japan, and the United States; zero telemetry; edge AI inference on every node. KOTO is the flagship workload.
- **Applied AI** — machine-learning models that operate seamlessly across distributed networks: autonomous synthesis, formally verified engineering, microsecond edge inference.
- **Quantum computing** — research systems, post-quantum migration planning, and hybrid classical–quantum architectures for durable enterprise capability.

**Leadership**

| | |
|---|---|
| Nicolai Kazuyuki Diaz Tsuzuki | Co-Founder & President (代表取締役社長) |
| Parker Czuba | Co-Founder & Chief Technology Director (取締役最高技術責任者) |

Winnipeg, Canada · Tokyo, Japan — [mutowa.com](https://mutowa.com)

*The philosophy: precision, longevity, and discipline — infrastructure designed to endure.*

---

*KOTO · Technical Documentation · © 2026 MUTOWA. All specifications subject to change as the network evolves.*
