Portable phase checkpoints¶
PalaestrAI checkpoints capable agents by default. A checkpoint is a complete, validated recovery point for one run instance and phase. PalaestrAI keeps only the latest complete checkpoint for that phase. Restarting is always explicit: a later run must ask to resume that latest checkpoint and name the original instance and phase.
This page explains the user-facing behavior. The runtime configuration is summarized in Runtime Configuration.
Normal run behavior¶
No checkpoint configuration is needed for the normal case:
data_path: /srv/palaestrai/eaton-demo
RuntimeConfig.data_path is resolved to an absolute directory. Checkpoint
packages live below its checkpoints subdirectory; do not copy, edit, move,
or select package files directly.
A package is created only when all of the following are true:
checkpointing is enabled;
every participating Brain supports portable checkpoints; and
the run reaches a completed evaluation boundary.
The boundary is quiescent: evaluation has finished and the agent contributes a consistent learning-state snapshot. It is not a mid-step or mid-episode save. Legacy Brains that do not provide the capability continue to run normally, but do not create a portable checkpoint.
PalaestrAI writes a new package before replacing the latest pointer. If a run is interrupted during a replacement, the previously completed checkpoint remains the resume candidate. If no completed evaluation boundary has been reached, there is nothing to resume.
Disable checkpointing¶
Set the single opt-out explicitly when a run must not create or resume portable checkpoints:
data_path: /srv/palaestrai/eaton-demo
checkpoint:
enabled: false
Do not combine enabled: false with a resume request.
Resume the latest checkpoint¶
Resume never happens automatically. To restart, use resume: latest and
identify the exact historical run instance and phase that own the latest
checkpoint:
data_path: /srv/palaestrai/eaton-demo
checkpoint:
enabled: true
resume: latest
resume_instance_uid: 4ed6c8bb-3d20-4ffb-b253-6c6ce0cb58a1
resume_phase: 0
resume_phase is the zero-based phase number. The selected package must
match the run, instance, phase, configured agent contracts, and configuration.
The new process performs setup first, restores the package before replay
pretraining or online learning, and then continues from the completed
evaluation boundary. PalaestrAI never scans other instances, chooses an older
package, or resumes a partial package.
Choose the same data_path that contains the original run’s checkpoint
packages and use the same Store. The Store is required for checkpoint evidence
and is checked during resume.
Verify a checkpoint or resume¶
For an evaluation-aware termination condition, the application log records an
accepted deterministic evaluation with its evaluation_id, checkpoint_id
and checkpoint_hash. A non-null checkpoint ID and hash identify the
committed package associated with that evaluation.
The Store contains a metadata-only CheckpointRef for each committed phase
package. It contains the checkpoint ID, package hash, relative package path,
run instance, phase, evaluation identity, completion time and retained state;
it never contains checkpoint binaries. Query retained references to confirm
the latest package for an instance and phase:
from palaestrai.store import Session
from palaestrai.store.query import checkpoint_refs
with Session() as session:
latest = checkpoint_refs(
session,
experiment_run_instance_uids=["4ed6c8bb-3d20-4ffb-b253-6c6ce0cb58a1"],
experiment_run_phases=[0],
retained=True,
)
print(latest[["package_hash", "relative_path", "completed_at", "evaluation_id"]])
After a successful resume, use the same log fields and Store query to confirm that later evaluation evidence refers to the restored run state and a newly committed latest package.
Troubleshooting¶
checkpoint.resume: latest requires ...Supply both
resume_instance_uidandresume_phase. PalaestrAI does not guess which historical run to restart.no latest checkpointor a missing latest packageThe selected instance and phase have no completed checkpoint in this
data_path. Check that the run reached an evaluation boundary, that its Brain supports portable checkpoints, and thatdata_pathpoints to the original output root.- Corruption, a manifest/hash error, or an incompatibility error
Stop the restart. PalaestrAI rejects incomplete, altered, or incompatible packages rather than loading them partially. Use the original environment, agent configuration, Store, and feature/action contracts, or start a new run without
resume: latest.- Insufficient free space
Free capacity on the filesystem containing
data_pathand rerun. Space for a complete replacement package is required; this safety margin is not a configuration option.worker: 1/ multi-worker checkpointing errorPortable checkpointing currently supports one rollout worker. Run with
worker: 1or disable checkpointing. A coordinated multi-worker quiescence protocol is required before multi-worker recovery can be safe.- Missing Store or CheckpointRef evidence error
Configure and retain access to the Store used by the original run. A checkpoint package is not published without its metadata-only Store evidence.
Safety and compatibility¶
PalaestrAI retains exactly one fully committed checkpoint per phase. The
current latest package is verified before it is resumed; package identity,
content hashes, Store evidence, run/phase identity, and agent/configuration
contracts must agree. Packages stay under the resolved data_path root and
are protected against unsafe paths and incomplete replacement. These are fixed
framework guarantees, not per-run tuning options.
The portable checkpoint records algorithm state and PalaestrAI’s run and condition state at a completed evaluation boundary. It does not claim to restore an interrupted environment step or a partially completed gradient update.