MUTOWA · Model Documentation · Current release: Super Heavy 2 · October 2026

KOTO

KOTO is our flagship frontier ensemble — a distributed network of 42 specialized 32B nodes coordinating in real time, powered entirely by solar energy across Canada, the United States, and Japan. The current lineup — KOTO Flash 1.2, KOTO Super 1.2, and KOTO Super Heavy 2 — runs on an ensemble that aggregates 1.34T parameters across its dedicated edge nodes for generating images, writing code, and making files.

Download this documentation (Markdown)


Contents

  1. Overview
  2. The Ensemble — how KOTO works
  3. Model tiers
  4. Capabilities and modalities
  5. Compute platform
  6. Network fabric
  7. Power and sustainability
  8. Security architecture
  9. Global footprint
  10. API quick start
  11. API reference
  12. Model behavior and identity
  13. Alignment and safety
  14. Release notes
  15. About MUTOWA

1. Overview

KOTO is MUTOWA's flagship frontier ensemble. Rather than a single monolithic network, KOTO distributes work across 42 specialized 32B nodes that coordinate in real time. Each node is sharpened for a domain — reasoning, code, vision, language, file construction — and an orchestration layer routes each request across the nodes best suited to it, so the ensemble behaves as one coherent model.

The result is a system that reasons deeply, responds fast, and produces finished artifacts: images, codebases, and files — end to end, from a single prompt.

MetricValue
Aggregate parameters1.34T across edge nodes
Context windowUp to 2,000,000 tokens (KOTO Super Heavy 2) · 1,000,000 (KOTO Super) · 128K (Flash)
Throughput~200 tokens per second (typical sustained, ensemble-wide)
Specialized nodes42
Availability target99.999%
Power source100% solar

KOTO runs on MUTOWA's Private Edge Compute network: solar-powered edge sites in Canada, the United States, and Japan, linked by millisecond anycast routing and post-quantum encrypted tunnels. Inference happens close to where requests originate — which is why KOTO is fast, private, and sustainable at the same time.

2. The Ensemble — how KOTO works

2.1 Specialized nodes

The ensemble is composed of 42 nodes, each a 32B-parameter specialist:

2.2 Real-time coordination

A request enters the ensemble once and is decomposed by a coordinator node into subtasks. Specialist nodes work in parallel where tasks are independent and in sequence where they depend on each other. Intermediate results pass between nodes as structured context, so the final answer is assembled — not averaged. This is a mixture-of-experts approach taken to the infrastructure level: expertise is a property of hardware placement, not just of model weights.

2.3 Extended chain-of-thought

On the thinking tiers, KOTO plans before it answers: it decomposes the problem, works through it step by step, and verifies its own conclusion against the question before responding. The thinking budget scales with the tier you choose.

2.4 Long context

KOTO Super Heavy 2 accepts up to 2,000,000 tokens of context (1,000,000 on Super) — whole codebases, full document corpora, or hour-long transcripts — with retrieval-augmented grounding across the window. Over-long conversations are compacted automatically rather than rejected: the oldest turns are summarized away while the latest request is always preserved intact.

2.5 Self-consistency sampling

For ambiguous, safety-critical, or multiple-choice questions, KOTO does not bet on a single generation. Following the self-consistency technique (Wang et al., 2022), it produces one deterministic answer plus several independent sampled attempts, and the majority answer wins. If no answer reaches consensus, the ensemble escalates automatically — more samples, a larger thinking budget, stricter answer extraction — and the caller simply receives a better-reasoned response. The mechanics are never exposed to the client.

3. Model tiers

KOTO is served through three tiers on the public API. All tiers share the same ensemble, weights, and knowledge — they differ in how much thinking and node coordination a request mobilizes.

Model idProfileThinkingContext windowOutput per runBest for
KOTO-flash-1.2 (alias KOTO-lite-1.2)Fast, directMinimal128K tokens8K tokensEveryday conversation, classification, extraction, high-throughput workloads
KOTO-super-1.2Balanced, flagshipExtended1M tokens1M tokensGeneral-purpose assistant work: writing, analysis, multi-step tasks, code
KOTO-super-heavy-2Deepest, maximum envelopeAlways on, maximum budget2M tokens2M tokensHard reasoning, agentic planning, whole-repo comprehension, system architecture, research-grade questions

Choosing a tier. Start with KOTO-super-1.2. Move down to KOTO-flash-1.2 when latency or volume dominates; move up to KOTO-super-heavy-2 when a task resists a first try — mathematics, long-horizon planning, subtle concurrency bugs — or when the input alone is a repository, a data dump, or a book.

Super Heavy 2 — the widest envelope in production. KOTO Super Heavy 2 carries a 2,000,000-token context window and a matching 2,000,000-token output budget per run — the only tier where what comes out can be as large as what goes in. Read an entire codebase, a legal discovery dump, or a book-length corpus in one prompt. Write a full novel, a complete platform, or an end-to-end build in one run. KOTO Super carries half: 1M in, 1M out. A single engine call is still bounded by the engine maximum; longer work is extended automatically across continuations until the run budget is reached, so the ceiling you can use is the one advertised. The legacy id KOTO-super-heavy-1.0 still resolves to the same tier.

Unknown or omitted model ids ride the flash tier — the API never errors on a model name.

4. Capabilities and modalities

4.1 Image generation

Prompt to pixels — any style, any scale. KOTO's vision nodes generate finished artwork end-to-end: no retouching, no external editor. Diffusion pipelines run directly on the ensemble's GPU accelerators.

4.2 Agentic coding

KOTO plans the architecture, writes the implementation, runs the tests, and fixes its own mistakes. One prompt in — a production-ready change out:

$ koto run "dashboard for womens world cup"

✓ Plan created — 4 tasks
✓ WebSocket server — Node.js, 42 lines
✓ Live metric cards — React 19
✓ Tests passed — 18/18
● deployed to preview in 38s

4.3 File construction

KOTO produces complete, verbatim multi-file artifacts — documents, data formats, structured exports — not sketches. Files are emitted with deterministic names and full contents.

4.4 Tool use and retrieval

KOTO can invoke tools (web search, retrieval) as part of answering, grounding real-time facts into its responses.

4.5 Modality support

ModalityInputOutput
TextYesYes
CodeYesYes
ImagesYesYes (diffusion nodes)
AudioYes—
FilesYesYes

5. Compute platform

Every KOTO node is a purpose-built edge computer, designed for silent, unattended, solar-powered operation.

5.1 Node hardware

ComponentSpecification
Compute nodes42
SoCAMD EPYC Embedded 3251 — 8 cores / 16 threads, 2.5 GHz base, 3.1 GHz boost
GPU acceleratorsIntegrated edge-GPU module + 1× high-end discrete accelerator
System RAM128 GB (4×32 GB) DDR5-5600 ECC unbuffered SO-DIMM per node
Ephemeral memory32 GB RAMDisk — tmpfs mounted noexec, nosuid, nodev
Boot drive1 TB Samsung 990 Pro PCIe Gen4 NVMe (OS image and kernel)
Thermal designPassively cooled, fanless aluminum chassis, 15 W–55 W configurable TDP
IsolationAMD SEV-SNP encrypted virtualization with PSP-managed memory keys

5.2 Why it looks like this

6. Network fabric

6.1 Capacity and latency

MetricValue
Total bandwidth16.8 Tbps
Average latency< 1.2 ms
Node latencyP99 round trip under 1.2 ms across regional fiber links
Network peeringTier-1
Uptime99.999%

6.2 Per-node networking

Packets are classified and steered at the driver level before the kernel wakes up — the ensemble's coordination traffic never contends with the OS network stack.

6.3 Long-haul links

Transpacific and transcontinental traffic rides dedicated fiber, including the FASTER and JUPITER submarine cable systems, with landings at Chikura and Shima (Japan) terminating to San Francisco and Los Angeles. Regional anycast routing keeps a request on the lowest-latency path between Canada, the United States, and Japan.

6.4 Protocols and SDKs

PlatformValue
ProtocolsgRPC, WebSocket
Official SDKsTypeScript, Go, Rust
Public API wire formatsStandard chat-completions and messages shapes

7. Power and sustainability

KOTO runs on 100% solar power. The ensemble's sovereign computational fabric is deployed across an isolated, dedicated 42-node bare-metal cluster operating on a private, microgrid-supported solar array.

7.1 The sovereign solar cluster

System attributeSpecificationOperational metric
Total compute nodes42 high-density multi-core heterogeneous nodes336 high-performance cores / 5.3 TB ECC RAM
Primary power sourcePhotovoltaic bifacial monocrystalline solar array145 kWp peak DC generating capacity
Energy storageLiFePO₄ (lithium iron phosphate) microgrid BESS720 kWh dedicated reserve (~18 h autonomy)
Cluster PUEPower Usage Effectiveness1.04 (direct negative-pressure ambient air heat exchange)
Lifecycle carbonScope 1 & 2 operational emissions0.00 g CO₂e / MegaToken (100% off-grid solar net-zero compute)
CoolingDirect convective air, negative-pressure flowSub-micron HEPA filtration, vertical thermal-stack exhaust
Water consumption—Zero gallons of potable water during inference cycles

7.2 Why direct air cooling

Hyperscale data centers frequently mask their true environmental footprint behind carbon-offset certificates while consuming clean regional drinking water for evaporative cooling. MUTOWA's array rejects this approach: ambient outside air is drawn through sub-micron HEPA filtration arrays, circulated directly over the 42 compute chassis, and vented vertically via thermal stack effects. The cluster consumes exactly zero gallons of potable water during inference cycles, and the realized facility PUE of 1.04 compares against the enterprise data center average of 1.58.

7.3 Solar-aware compute pacing (SALD)

The 42 nodes are managed by MUTOWA's proprietary Solar-Aware Load Director (SALD). Inference tasks are classified into two latency vectors:

7.4 Efficiency by design

The sustainability thesis is simple: clean power where the work happens. Millisecond compute does not require a mega-datacenter; it requires well-placed, efficient machines on renewable generation.

8. Security architecture

KOTO's security model is post-quantum by default and attested in hardware.

LayerMechanism
Encryption in transitNIST ML-KEM (Kyber-1024) key encapsulation, AES-256-GCM WireGuard tunnels
Perfect forward secrecySession keys rotate every 60 seconds via CPU true RNG
Encryption in useAMD SEV-SNP encrypted virtualization, PSP-managed memory keys
Ephemeral state32 GB RAMDisk, mounted noexec / nosuid / nodev
TelemetryZero — requests are not logged, profiled, or used for training

8.1 Post-quantum mesh

All inter-node and client-facing tunnels use Kyber-1024 key encapsulation layered under AES-256-GCM inside WireGuard. The mesh is therefore resistant to harvest-now-decrypt-later collection: traffic captured today remains unreadable to quantum adversaries later.

8.2 Sixty-second key rotation

Every tunnel's session key is re-derived from the CPU hardware random number generator every 60 seconds. A compromised key exposes at most one minute of a single session.

8.3 Attested enclaves

Node workloads execute inside AMD SEV-SNP encrypted virtual machines with measured boot. Memory is encrypted with PSP-managed keys that never leave the silicon; remote attestation proves the exact code running before a node joins the ensemble.

8.4 Privacy posture

MUTOWA operates Private Edge Compute with zero telemetry: KOTO does not retain your prompts, outputs, or usage fingerprints for its own purposes. Inference state lives in the noexec RAMDisk and evaporates on session end.

9. Global footprint

RegionNodes
Canada14
United States20
Japan8
Total42
CoverageValue
Live cities7
Countries3

The footprint follows MUTOWA's founding geography — Winnipeg, Canada and Tokyo, Japan — with dense United States coverage for transpacific peering. Anycast routing directs each request to the healthiest, nearest capacity; regional failures drain to neighboring sites automatically, which is where the 99.999% uptime figure comes from.

10. API quick start

The KOTO API is served at https://api.mutowa.com. Endpoints are drop-in compatible with the two most common chat-API wire shapes, so most existing tooling connects by pointing it at the KOTO base URL with a KOTO key.

10.1 List models

curl https://api.mutowa.com/v1/models

10.2 Chat-completions endpoint (curl)

curl https://api.mutowa.com/v1/chat/completions \
  -H "Authorization: Bearer mk_…" \
  -H "content-type: application/json" \
  -d '{
    "model": "KOTO-super-1.2",
    "messages": [
      {"role": "system", "content": "You are a precise research assistant."},
      {"role": "user", "content": "Summarize the state of fault-tolerant quantum error correction."}
    ]
  }'

10.3 Messages endpoint (curl)

curl https://api.mutowa.com/v1/messages \
  -H "x-api-key: mk_…" \
  -H "content-type: application/json" \
  -d '{
    "model": "KOTO-super-heavy-1.0",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "Design a rate limiter for a multi-region API."}]
  }'

10.4 Streaming (SSE)

Set "stream": true on either wire format. Deltas arrive as server-sent events; streams are unbuffered end to end, so first-token latency reflects ensemble dispatch rather than proxy buffering.

10.5 Authentication

All completion endpoints require a Bearer key — Authorization: Bearer mk_… (the messages shape also accepts x-api-key). Keys are issued by MUTOWA and can carry per-key rate limits (requests per minute, tokens per minute) and output caps.

11. API reference

11.1 Endpoints

EndpointShapeDescription
GET /v1/models—List the KOTO tiers
GET /api/health—Service health and dependency status
GET /docs—This documentation (HTML)
GET /docs/download—This documentation (Markdown download)
POST /v1/chat/completionsChatChat completions; supports stream, temperature, top_p, max_tokens
POST /v1/completionsChatLegacy text completions
POST /v1/messagesMessagesMessages API; supports stream, max_tokens, system
SDK base-URL alias of POST /v1/messagesMessagesAlias route for SDKs that pin a fixed base URL (exact spelling in the GET /v1 index)

11.2 Request parameters

ParameterTypeDefaultDescription
modelstringKOTO-flash-1.2Tier id. Unknown or omitted ids ride the flash tier — never an error. The retired KOTO-lite-1.2 and -contributor spellings are accepted as aliases.
messagesarrayrequiredConversation turns. Over-long conversations are compacted automatically rather than rejected.
temperaturenumber0Sampling temperature. The API defaults to deterministic output for reliability.
top_pnumber—Nucleus sampling.
max_tokensintegerplatform defaultOutput cap per run: tier ceilings lite 8K · super 1M · super-heavy 2M (per-key caps only lower). Thinking headroom is added on top automatically on thinking tiers; a single engine call is clamped to the engine maximum and longer answers extend across automatic continuations.
streambooleanfalseStream deltas over SSE.

11.3 Errors

Errors return structured JSON — never internal traces.

StatusMeaning
400Malformed request (missing messages, invalid parameter)
401Invalid or missing API key
429Per-key rate limit or token budget exceeded (JSON body gives details)
502Service momentarily saturated — safe to retry

11.4 Operational guardrails

Background jobs and agent loops must never run away. KOTO enforces hard ceilings across all client and server boundaries; exceeding one surfaces as a structured error or a clean truncation — never a hang or an internal trace. The values below are platform defaults for this release and are enforced server-side:

SubsystemHard limit boundaryEnforcement mechanism
Token burst ingress20,000 tokens per minuteToken-bucket rate limiter; queues requests exceeding ceiling
Active context bufferStrictly 12 previous dialogue turnsWindowed truncation; prevents quadratic prompt inflation
Agent loop executionMaximum 6 tool rounds per turnHalts execution chain and presents collected findings
Web search operationsMaximum 2 rounds per query turnPrevents infinite search loops on ambiguous queries
Subprocess code sandbox30-second timeout / 5 MB output limitImmediate SIGKILL termination on ceiling breach
Local file I/O≤ 1 MB per atomic file operationSafePath validator halts read/write on oversized payloads
Browser extraction≤ 100 KB raw DOM text (32 KB default)Triggers clean Markdown extraction and truncates
ZIP compilation≤ 50 files / ≤ 5 MB total payloadIn-memory buffer enforcement prior to disk stream
Mesh complexity≤ 20,000 total trianglesProcedural geometry-index ceiling to prevent GPU crash

12. Model behavior and identity

KOTO identifies as KOTO, an AI assistant made by MUTOWA. Its behavior contract:

About MUTOWA — from 無永遠 (mu eien), "boundless permanence" — MUTOWA is a technology company building enduring, fault-tolerant infrastructure across quantum-ready systems, applied AI (including KOTO), and private edge computing. Founded by Parker Czuba (Co-Founder & Chief Technology Director) and Nicolai Kazuyuki Diaz Tsuzuki (Co-Founder & President), operating from Winnipeg, Canada and Tokyo, Japan. Website: mutowa.com.

13. Alignment and safety

KOTO is trained with reinforcement learning from human feedback and a constitutional training regime, evaluated continuously for factuality, instruction fidelity, and refusal calibration. Alignment research runs in the loop across releases rather than being applied after the fact.

Toward the mission: every KOTO release narrows the gap between narrow tools and general intelligence — deeper reasoning, longer memory, richer action — safely, and on purpose.

14. Release notes

AreaCurrent lineup — Flash 1.2 · Super 1.2 · Super Heavy 2
Ensemble42-node real-time coordination; 1.34T aggregate parameters across dedicated edge nodes
ReasoningExtended chain-of-thought with tier-dependent thinking budgets
VisionDedicated diffusion edge nodes for end-to-end image generation
Agentic codingPlan → implement → test → repair loop producing production-ready changes
ContextTiered context windows — 2M tokens on Super Heavy 2, 1M on Super — with automatic conversation compaction
OutputPer-run output budgets to match the window: 2M tokens on Super Heavy 2, 1M on Super, delivered across automatic continuations
APIThree public tiers — KOTO-flash-1.2, KOTO-super-1.2, KOTO-super-heavy-2 (legacy KOTO-lite-1.2 and KOTO-super-heavy-1.0 aliased) — on standard chat-completions and messages endpoints
InfrastructureFully solar-powered footprint; post-quantum mesh (Kyber-1024) with 60-second key rotation; 99.999% uptime

Roadmap

15. About MUTOWA

MUTOWA — from the Japanese 無永遠, boundless permanence — is a technology company linking quantum computing, applied artificial intelligence, and private edge computing: building foundational frameworks tailored for the next generation of computation.

Three connected practices

Leadership

NameRole
Nicolai Kazuyuki Diaz TsuzukiCo-Founder & President (代表取締役社長)
Parker CzubaCo-Founder & Chief Technology Director (取締役最高技術責任者)

Winnipeg, Canada · Tokyo, Japan — mutowa.com

The philosophy: precision, longevity, and discipline — infrastructure designed to endure.


KOTO · Technical Documentation · © 2026 MUTOWA. All specifications subject to change as the network evolves.