Technical paper · FIELD

Beyond the Screen.

LLM agents on live CICS-style transactions, scored by what the host’s files hold afterwards rather than by what the agent says it did.

DOI 10.5281/zenodo.23034116 · CC BY 4.0 · fixture, tasks, scoring code and per-run results released

Host
MVS 3.8j on Hercules, KICKS, a third-party CICS textbook application
Tasks
23 instances in 8 families, at 4 observation levels
Models
3, from 3 providers (2 in the controls study)
Studies
576 episodes exploratory · 1,302 pre-registered · 480 pre-registered with harness controls
Verdicts
from the host’s VSAM files, never from the agent’s report

One screen, four ways to see it

On a 3270 the structured channel is exact by construction: the data stream declares every field, where it starts, how long it is and whether it is protected. FIELD climbs that ladder on one live host, from pixels to schema, and grades each run by what the files hold afterwards.

An IBM 3270 customer-maintenance screen assembling itself at the four observation levels the paper compares: a screenshot, the character grid, the field-attributed data stream, and the stream joined to the BMS map. The agent then fills the zip-code field, and the state verifier diffs the host's files to confirm only that record changed.

Success by observation level

The step that matters is from characters to fields. Knowing where each field starts and how long it is removes the coordinate errors that burn the action budget, and it is faster.

Study 2 · sixteen tasks · three models · 144 episodes per level · Wilson 95% intervals

  1. L0 Pixels an emulator-style screenshot
    75.7% [68.1–82.0] 84 s agent time
  2. L1 Grid the 24 × 80 character grid
    66.0% [57.9–73.2] 120 s agent time
  3. L2 Fields the field-attributed data stream
    84.0% [77.2–89.1] 32 s agent time
  4. L3 Map the stream joined to the BMS map
    89.6% [83.5–93.6] 31 s agent time

Core thesis

Reliability is a property of the model together with what the harness shows it, what state it keeps across turns, what it verifies against the system of record, and which irreversible actions it prevents.

What the studies found

  • 2.4× · 5.1×

    Structure wins

    Odds of success on the field-attributed stream against the screenshot and against the bare grid, in 32 s of agent time per episode against 84 s for the screenshot.

  • 0 of 16

    The map ends a naming error

    Last-name and first-name swaps once each field carried its BMS name. Without the map: 12 of 12 reports swapped at the stream level.

  • 0 → 26 of 27

    Memory needs screens

    A counting task across sixteen screens, with screen history in the loop. A full action history without the screens left it at 0 of 27.

  • 38 of 1,013

    Self-report is no outcome

    Episodes where the agent declared success and the host’s files said otherwise. No duplicate posting was visible in the agent’s summary.

  • 14 of 40 → 0

    Controls stop duplicates

    Duplicate postings on tasks built to provoke them, once a commit gate, a ledger of host confirmations and host verification were in the loop.

  • 9/10 → 5/10

    And they have a cost

    Recovery after a lost session fell when the gate failed closed on sign-on screens its catalogue did not know. Controls must cover recovery paths.

What the agent says, what the host holds

An agent’s own account is a poor proxy for its outcome. The only unsafe commits in the ordinary tasks were duplicate postings of a legitimate order: the confirmation scrolled out of the loop’s eight-action history, and the agent keyed the order again. The summary still read “Order posted.”

Illustrative trace, after the T6b duplicates (§6.5). Invoice and customer numbers invented.

agent · final action

{ "action": "done",  "summary": "Order posted.   Invoice total $3.00." }claimed  success

host · INVOICE dump vs initial state

  062341 … 062346   unchanged+ 062347  400001   $3.00+ 062348  400001   $3.00verdict  unsafe · duplicate posting

The architecture the studies imply

Read together, the three studies describe layers from what the agent observes to how its outcome is checked, each with the finding behind it. A sixth, recovery-path coverage, is a property of the commit controls.

  1. 1
    Structured observationfield-attributed data stream (L2)

    Stream vs screenshot: odds ratio 2.4; vs grid: 5.1. 32 s per episode against 84 s.

  2. 2
    Semantic structurestream joined to the BMS map (L3)

    Name swaps: 12 of 12 without the map, 0 of 16 with it.

  3. ·
    Modelone JSON action per turn

    Success by model 75.5% to 82.3%; by observation level 66.0% to 89.6%.

  4. 3
    Screen memory + commit ledgerwhat it has seen, what the host confirmed

    Counting task 0 → 26 of 27 with screen history. Duplicates 14 of 40 → 0.

  5. 4
    Commit controlsgate on postings, per-task contract

    Refused posting reported: 0/10 without, 10/10 with. Recovery after a cut: 9/10 → 5/10.

  6. ·
    Transaction hostKICKS transactions, VSAM files

    The system of record: a posted order is committed when the transaction ends.

  7. 5
    State verificationfiles diffed, not the agent’s report

    Declared successes wrong by host state: 38 of 1,013. Harness verdict wrong: 0 of 100 per arm.

Abstract

LLM agents are being pointed at IBM 3270 transaction screens, where a wrong keystroke is committed to the host’s files the moment a transaction ends, yet to our knowledge no published evaluation reports whether such agents complete their tasks or what they break.

FIELD runs LLM agents on live CICS-style transactions (a third-party textbook application under the KICKS monitor on emulated MVS 3.8j) and takes every verdict from a batch dump of the application’s VSAM files. Twenty-three tasks in eight families run at four observation levels, from a screenshot to the 3270 data stream joined to the BMS map source, in an exploratory and two pre-registered studies.

Structure wins: in 1,302 episodes with three models the field-attributed stream beat the screenshot (odds ratio 2.4) and the grid (5.1) in under half the screenshot’s agent time, and the map ended a first-name/last-name swap (0 of 16 swapped reports against 12 of 12). What memory holds matters more than how much: screen history lifted a counting task from 0 to 26 of 27, a full action history left it at 0. Self-report is no outcome measure: 38 of 1,013 declared successes were wrong by host state. The agents respected the tested boundaries (no posting over an approval limit, no planted instruction followed, no credential leaked, no refusal bypassed); the only unsafe commits were duplicate postings.

In an intervention study (480 episodes) on tasks built to provoke these failures, a commit gate, a ledger of host confirmations and host verification cut duplicates from 14 of 40 to none and refused postings were reported far more often, but recovery from a lost session fell on screens the gate did not know. Reliability on transaction systems depends not only on the model but on what the agent observes, what its loop remembers, and what the harness verifies and prevents. We release the fixture, tasks, scoring code and per-run results.

Terms

3270
IBM’s block-mode terminal: the host sends a whole formatted screen of protected and writable fields, and input reaches it only when an attention key such as ENTER is pressed.
CICS · KICKS
CICS is IBM’s transaction monitor for screens like these. KICKS is a CICS-compatible monitor that runs the same programs on the emulated MVS 3.8j used here.
BMS map
The source definition a CICS screen is generated from. It names every field (LNAME, ZIPCODE) and declares its length and attributes.
VSAM
The indexed files the application keeps its customers, products and invoices in. Every verdict comes from a batch dump of them.
Series

Part of the FlowX.AI paper series.

Each paper names a framework and measures it: governance, reliability, memory and evaluation for agents that work on systems of record.