We present CA-WM, a world model for grid-based agents in which an LLM induces action-conditional transition rules over $3\times3$ local neighborhoods, stored in a persistent rule bank. Locality and determinism are enforced by construction, yielding position-independent rules that transfer to unseen layouts where covered; completeness is validated after every LLM call. The correction mechanism is the core contribution: when a predicted transition mismatches the observed outcome, the wrong rule is evicted and re-learned from a single example within the same step. This allows CA-WM to override an incorrect prior in the tested setting --- on ForwardDoorKey, a MiniGrid variant where \textsc{forward} opens a locked door, CA-WM achieves 100\% planning success on five out-of-distribution layouts against 0\% for a physics oracle that cannot update its assumptions. On standard DoorKey-5$\times$5 the bank saturates after one accumulation episode: subsequent episodes require zero LLM calls while exactly predicting every changed cell --- a $\approx\!25{,}000\times$ reduction in marginal prediction cost. A greedy lookahead planner over the bank reaches BFS-optimal solutions on all out-of-distribution seeds. All results use \texttt{Qwen3-Next-80B-A3B-Instruct} (3B active per token) served locally via vLLM at zero API cost.

CA-WM: Learning Cellular Automaton World Models via LLM Rule Induction and Online Correction

Cesare Zavattari;Alessandro Tommasi;Giuseppe Prencipe
2026-01-01

Abstract

We present CA-WM, a world model for grid-based agents in which an LLM induces action-conditional transition rules over $3\times3$ local neighborhoods, stored in a persistent rule bank. Locality and determinism are enforced by construction, yielding position-independent rules that transfer to unseen layouts where covered; completeness is validated after every LLM call. The correction mechanism is the core contribution: when a predicted transition mismatches the observed outcome, the wrong rule is evicted and re-learned from a single example within the same step. This allows CA-WM to override an incorrect prior in the tested setting --- on ForwardDoorKey, a MiniGrid variant where \textsc{forward} opens a locked door, CA-WM achieves 100\% planning success on five out-of-distribution layouts against 0\% for a physics oracle that cannot update its assumptions. On standard DoorKey-5$\times$5 the bank saturates after one accumulation episode: subsequent episodes require zero LLM calls while exactly predicting every changed cell --- a $\approx\!25{,}000\times$ reduction in marginal prediction cost. A greedy lookahead planner over the bank reaches BFS-optimal solutions on all out-of-distribution seeds. All results use \texttt{Qwen3-Next-80B-A3B-Instruct} (3B active per token) served locally via vLLM at zero API cost.
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11568/1369429
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact