movingparticle@caja-negra RODRIGO POLO EC, GYE ~/reto-2/DESIGN_DOC --:--:-- ES

movingparticle@caja-negra:~$ cat ./SYSTEM_DESIGN.txt

FRONTIER Bogotá 2026 · Challenge 2 · 12:00 milestone.
A service has been running in production for 10 years and nobody knows how it works. We'll discover its behavior by observing it and build a replica that responds exactly the same — errors included — without ever calling it at runtime.

01 // what

Replace a system with no code and no documentation with a deterministic, byte-for-byte replica, against cases we don't know.

02 // how

Hypothesis-driven probing to infer the real rule, and a native in-memory HTTP server that emulates commands, states and errors.

03 // proof

Differential comparison oracle vs. replica, an adversarial battery, and exact parity of status and body.

status+body fidelity active learning node esm · 0 deps atomic reset no overfitting
documentation architecture ↗
overview

Summary

A 30-second read. The technical detail is further down, in the same document.

The problem

The inside is unknown. We only see INPUT → BOX → OUTPUT. We must reconstruct the rule, not memorize examples.

The solution

A discovery harness that queries the original, and an HTTP entrypoint (Cloudflare Worker + Durable Object) that responds the same way.

The evidence

Every behavior cites a line of the log. Every response is compared against the original. Nothing is claimed without proof.

PRINCIPLE. A function that reproduces the known examples does not necessarily reproduce the real rule. We design to generalize, not to pass the visible set.

Critical points

critical pointstatusresolution
Replica not implementedresolved6 commands + 17 errors in src/engine.mjs; parity 20/20.
Unknown control algorithmresolvedISO 7064 Mod 97-10 (16/16 cases + edge probes).
State not isolated / multi-instanceresolvedSingle Durable Object global-caja-negra (one state).
Entrypoint not deployedresolvedCloudflare Workers: https://caja-negra-reemplazo.parsec-ai-labs.workers.dev.
15:30 incidentopenFrozen baseline + adaptation with no regression (INCIDENT.md).
Contract notes (assumptions). GET /health requires a token: without one → 401 E001 (the harness sends a token on every call). The body of /__reset is {"ok":true} and is not compared by the harness. Both remain assumptions until verified against the original.
observed data

Stats

Two sources side by side: green = 20 public examples (pruebas/ejemplos.jsonl) · cyan = 14 live probes against the oracle (experimentos.jsonl). Block bars, sorted, with visible n. No pies.

20
examples
13
2xx responses
7
errors
6
known commands
1
business route

http status baseline · 20 examples

n = 20 · sorted by frequency

200 OK
11
400 BAD
4
201 NEW
2
422 LIM
2
404 MISS
1

http status probes run · 14 probes

n = 14 · live adversarial probes

400 BAD
9
404 MISS
3
200 OK
2

error families baseline · 20 examples

n = 7 · E2xx shape · E3xx command · E4xx resource · E5xx business

E2xx
2
E3xx
2
E5xx
2
E4xx
1

error families probes run · 12 errors

n = 12 · E1xx appears (message shape), which the baseline never reached

E3xx
6
E4xx
3
E1xx
2
E2xx
1

commands baseline · 20 examples

n = 20 · the command is case-insensitive (CONSULTa was accepted)

CONSULTA
5
ALTA
4
TRANSFER
4
DEPOSITO
3
ANULA
2
PING
1
UNKNOWN
1

commands probes run · 14 probes

n = 14 · adversarial bias: TRANSFER and ANULA to force edges

TRANSFER
4
ANULA
3
DEPOSITO
2
PING
1
ALTA
1
CONSULTA
1
RETIRO
1
(empty)
1
BASELINE (20)          n    bar          PRUEBAS (14)        n    bar
200 OK                11   ███████████  400 BAD             9   █████████
400 BAD                4   ████         404 MISS            3   ███
201 NEW                2   ██           200 OK              2   ██
422 LIM                2   ██
404 MISS               1   █
Green = baseline (20 public examples) · cyan = probes run (14 live probes, ids 21–34 in experimentos.jsonl). Probing settled error precedence (E303→E304→E401→E506), amount limits and the control algorithm (mod 97 checksum).
system design · understanding

Problem

arc42 §1–3 and the rubric's understanding criterion: goal, boundaries, assumptions, what is out of scope and why.

Goal

An autonomous, deterministic, public replica that responds the same as the original — same status and bodies — against a hidden set of 100 tests.

Inputs / Outputs

  • In: POST /msg with {"m":"CMD;arg;...;ctrl"} + Bearer token.
  • Out: exact HTTP status + exact JSON or text body.
  • Control: GET /health, POST /__reset.

Explicit constraints

  • The replacement cannot call the original or another replacement.
  • 5 s per step; 10 req/s and 8,000/day per token.
  • Public repo with no private organizer material.
  • Own fuzzers: allowed and expected.

Implicit constraints

  • A single process implies per-instance state.
  • The service may change during the day (15:30 incident).
  • With no token or a foreign token, behave like the original.

Assumptions

  • Deterministic state machine.
  • Positional arguments separated by ;.
  • The last field validates message integrity.
  • Ephemeral state: /__reset goes back to zero.

Out of scope

  • Persistence: the contract asks for reset, not durability.
  • Proxy to the original: the rule scores it 0.
  • Runtime UI: it adds no parity.

Edge cases

empty bodycase-insensitive E302 arityE201 control E202 controlE401 account E506 limitE507 balance already voidedunknown route large payloadinvalid encoding

Failure modes

  • Overfitting to 20 examples, fails the new ones.
  • Validation order mis-inferred, wrong error.
  • Shared state without reset, cascade.
  • Schema change from the incident, broken responses.
  • Latency > 5 s, timeout = failure.
system design · architecture

Architecture

Production on Cloudflare Workers with a Durable Object that guarantees a single global state; the same engine runs on a local Node server. Integrity with ISO 7064 Mod 97-10.

Public entrypoint: https://caja-negra-reemplazo.parsec-ai-labs.workers.dev

→ open architecture.html (C4 level 1 and 2)

trade-offs · madr 4.0

Decisions

Each ADR states what was discarded, why, what is gained, what is sacrificed and when it would change.

ACCEPTED ADR-0001 · Node.js ESM, zero dependencies

contextReproducibility and startup on any evaluation host.
driversReproducibility · attack surface · startup time
options(A) Fastify/Express + TypeScript + Zod · (B) native node:http
We chose (B). With no dependency tree, deployment is identical on any machine.
+ determinism, portability, startup < 50 ms
- routing and validation written by hand
change if> 25,000 req/s with dynamic schemas: Fastify.
confirmationnpm ci && npm test && npm start in a clean environment.

ACCEPTED ADR-0002 · Hypothesis-based probing

contextShort budget: 10 req/s, 8,000/day.
driversInformation per probe · SPEC/LOG traceability
options(A) random fuzzing · (B) probes that separate hypotheses
We chose (B). Each probe tries to falsify a hypothesis and lands in the LOG.
+ semantics, citable evidence, anti-overfitting
- slower for purely syntactic junk
change ifFree-form fields or deep JSON appear: add mutators.

ACCEPTED ADR-0003 · Single state with Durable Object + atomic reset

contextThe judge calls /__reset before each of 100 tests and assumes a single process/state.
driversSingle global state · isolation between tests · reset latency
options(A) classic SQLite/Redis · (B) stateless instances · (C) Durable Object with a fixed name
We chose (C): a global-caja-negra Durable Object guarantees a single logical state, with an in-memory reset (< 1 ms) and no disk.
+ single state across instances, instant reset, trivial deploy
- state lives in the DO and requires a binding (it is not a plain Node process)
change ifCrash-recovery or multi-region shared state were required.

ACCEPTED ADR-0004 · Adapter for the incident

contextAt 15:30 the service may change format.
options(A) rewrite the engine · (B) normalization adapter
We chose (B). The change lives in a single point. What already worked keeps working.
+ minimal patch, frozen regression
- one more indirection
change ifThe incident invalidates the domain model, not just the format.
plan to measure

Measurement

How we'll know it works before building it. No number, no defense.

Baseline

After /__reset, the 20 examples give 20/20 on status and body (automated test).

Metrics

  • Status + body parity
  • Step < 5 s
  • Reset < 1 ms
  • Suite 11/11 green

Success

  • 100% on the known set
  • 0 post-incident regressions
  • Isolated reset
  • Valid entrypoint contract
testwhat it tries to breakexpectedprio
happy pathPING → ALTA → DEPOSITO → CONSULTAexact status and balancesP0
empty/msg with no bodysame error as the originalP0
unknownSALDOS;...400 E301P0
aritytoo many or too few ;400 E302P0
controlnon-numeric field400 E201 / E202P0
accountCONSULTA;AC-9999404 E401P0
limittransfer over the threshold422 E506P1
balancetransferring more than available422 E507P1
void x2ANULA of an already voided operationok:false, ya anuladaP1
bad routeGET /msg, unknown pathsame 404/405P1
payloadlarge body or broken encodingrejection with the exact codeP1
reset/__reset without a tokenlike the original, without leaking the mechanismP0
incidentthe case that changes at 15:30adapted, with no regressionP0
latencyeach step< 5 sP2
hidden tests

Hypotheses

confirmed repeats · likely seen once or inferred · assumption is the replacement's default.

idbehaviorevidenceconfidence
C1PING → {pong:true} with no state#1confirmed
C2command is case-insensitive#11 CONSULTaconfirmed
C3ALTA creates an incremental AC-#####3 #4confirmed
C4the control is validated before acting#2 #17likely
C5DEPOSITO credits and returns operacionId#6 #18 #19confirmed
C6TRANSFER moves balance#7 #8confirmed
C7there is a cap E506 and balance E507#9 #16confirmed
C8ANULA is idempotent#13 #14confirmed
C9unknown command → literal 400 E301#12confirmed
C10invalid arity → 400 E302#20confirmed
C11control = ISO 7064 Mod 97-10 (A=1…Z=26, no padding, CD = 98 − (N·100) mod 97)16/16 + probesconfirmed
H1the control is a check digit of the payloadconfirmed as ISO 7064confirmed
H2order: shape → control → resource → business#15 vs #9likely
H3IDs are per process, they reset with reset#3 #6likely
S1large payload or invalid encoding → stable errorcontractassumption
S2unknown route → same 404contractassumption
Anti-overfitting: for every confirmed rule, a probe where two hypotheses predict differently. The one with the highest information gain is run.
implementation

Plan

order

  • P0 server: /health /__reset /*
  • P0 engine + store with atomic reset
  • P0 guided probing (experimentos.jsonl)
  • P1 adversarial battery
  • P1 15:30 incident: detect, isolate, patch, regression
  • P2 README + 2 min video

open

  • Resolved: control = ISO 7064 Mod 97-10.
  • Resolved: precedence E303 > E304 > E401 > E506 > E507.
  • What exactly changes at 15:30?
  • Does the rate-limit return a body we must emulate?
  • Do alternate routes (e.g. /saldo) exist in the original?
debrief

Defense

Pattern: we chose X for A and B; we discarded Y for C; the trade-off is D; if E changes, we move to Y.

Why this architecture?
It separates offline discovery from online execution. The replacement never calls the original.
What did you reject?
A proxy to the original (the set scores 0) and a blind fuzzer (it doesn't model state).
The weakest part?
The exact precedence between shape errors (a couple of extra probes were needed) and the single state, which depended on using a Durable Object.
What breaks at 10x?
Several instances, each with its own RAM. The next step would be shared state with a lease.
How do you know it works?
Diff oracle vs replica on the 20 examples and on the adversarial battery. Same status, same body.
What do you assume?
Deterministic state machine, positional commands, resettable state.
What if a dependency goes down?
The runtime has no dependencies. If the oracle goes down, discovery stops; the replica keeps going with what was already inferred.
What did the AI do and what did you change?
The AI proposed structure and candidate formulas. We accepted or dropped them with probes. The generated validation order was corrected.
With more time?
1) close the control. 2) multi-instance state. 3) health and operation metrics.
Why does it generalize?
We model the rule, not the sequence of examples. Every behavior was falsified against a rival hypothesis.
keys 1234 5678 jump to section