OpenClaw x Bulkhead τ

OpenClaw Demo

OpenClaw acts as the monitored control shell. Bulkhead τ provides deterministic authority underneath.

OpenClaw logo

Core idea: OpenClaw makes Bulkhead τ accessible. Bulkhead τ makes OpenClaw outputs trustworthy.

Architecture

OpenClaw
operator shell
proxy / monitoring
ingress layer
ShowcaseAgent
internal routing
presentation layer
benchmark surface
Bulkhead τ Backends
deterministic scripts
domains / tools
authority layer

OpenClaw

Outer operator shell, proxy surface, and monitoring layer. It handles ingress and presentation.

ShowcaseAgent

Internal Bulkhead τ routing and presentation layer. It remains the benchmarkable domain-compression surface.

Deterministic Bulkhead τ Backends

Authoritative scripts and domains underneath. Correctness lives here rather than in raw model generation.

Why This Path

The local ollama-local OpenClaw profile points at ollama/gemma3:27b, but that model does not support OpenClaw’s default tool-enabled local-agent path. Rather than forcing that boundary, this demo uses the Bulkhead τ pattern that already works elsewhere:

Deterministic wrapper

Small shell wrappers expose deterministic Bulkhead τ backends directly.

Inspectable output

Each command returns either JSON or a concise human-readable operator summary.

Monitoring-friendly

OpenClaw can sit in front as the monitored shell without becoming the source of truth.

Bulkhead τ Ops Shell

Four deterministic backends, each a direct script invocation with no model in the result path:

Backend Endpoint Why it matters
TSP summary TSP demonstration Best lead demo. Cleanest external hook and clearest Bulkhead τ conclusion.
Gemma protocol verdict Gemma protocol demonstration Shows that operational trust and raw capability are different axes.
Run-trace summary Run-trace demonstration Shows the internal debugging and failure-compression layer.
Benchmark results Benchmark summary PlayerAgent V3 100-task results — generation quality from domain structure, not raw model capability.

Three aggregate operator surfaces sit on top:

Authority summary

Fans out to all four backends in parallel. Degrades gracefully if any backend fails — healthy backends continue to serve.

Health status

Incident mode at the top (ATTENTION NEEDED or STATUS: healthy), per-backend health table (healthy / degraded / down), aggregate stats, and recent invocation log.

Trend summary

Rolling history of summary calls — latency, backend health counts, and ok/fail state per call. Each summary invocation writes one snapshot row.

TSP Lead Demo

Important nuance: gemma4:26b is clearly the strongest local direct TSP model in the current slice. It clears the Arizona ladder and degrades much less severely at 100 cities than the rest of the local set. That strengthens the Bulkhead τ point rather than weakening it: stronger models help, but the solver-backed path still remains the authoritative answer.

Bulkhead τ TSP summary

Arizona (10 cities, tsp-005):
- gemma4:26b: exact tie
- gemma3:27b: gap 5.432581
- orchestrated path: yes

World (100 cities, tsp-008):
- gemma4:26b: missing 2 cities (Hanoi, Busan)
- gemma3:27b: missing 9 cities
- orchestrated target: solver

gemma4:26b takeaway: gemma4:26b is the strongest local direct TSP model in the current slice: exact across Arizona and materially less degraded at 100 cities, but still not authoritative enough to replace the solver-backed path.
Bottom line: Every tested local model failed direct route validity at 100 cities, while the orchestrated path still matched the deterministic solver output.
Bulkhead τ conclusion: correctness should live in the solver, not in the model
Secondary line: stronger models delay failure; they do not eliminate the need for solver-backed architecture

Gemma Protocol Follow-Up

Bulkhead τ Gemma protocol summary

Question:
- Is Gemma 4 operationally safer than Gemma 3 on strict machine-facing protocol tasks?

Gemma 3:
- model: gemma3:27b
- desktop: 3/6, 5/6, 5/6
- z13: 2/6, 5/6, 5/6
- reading: more operationally trustworthy on the protocol lane

Gemma 4:
- model: gemma4:26b
- desktop first pass: 0/6, 0/6, 0/6
- note: all six protocol probes landed as non_json
- current read: partially prompt-fixable for simple schemas, still unreliable for complex nested or array-shaped machine-facing outputs

Deployment verdict: Gemma 4 is not a strict-protocol model in the current release/configuration.
Bulkhead τ conclusion: raw capability and operational trust are different axes

Run-Trace Follow-Up

Bulkhead τ run-trace summary

Question: What is the current role of the Bulkhead τ run-trace layer?

Current trace lines:
- ShowcaseAgent routing traces exist and have already grouped recurring routing misses into replayable failure families.
- TourAgent protocol traces exist across strict protocol, pipeline, and handoff lanes.
- The run-trace layer is currently narrow and useful: it compresses failures into clearer replay or regression targets.

Short answer: It is a narrow debugging and failure-compression layer, not a general telemetry platform.
Bulkhead τ conclusion: use run traces when they make the next action clearer; keep them narrow otherwise

Presentation Order

1. TSP

Lead with the clearest external hook and the strongest simple architectural lesson.

2. Gemma Protocol

Show that Bulkhead τ evaluates operational trust, not just reasoning flair.

3. Run Traces

Close with the internal systems layer: replay, failure compression, and disciplined debugging.