# CRDT Snapshot Compaction ## Problem The huskies project CRDT has grown to 55K ops / 276MB in `pipeline.db`. Every container restart replays all operations in a tight synchronous loop on the tokio runtime (`crdt_state/state/init.rs:68-78`), taking 27+ minutes and freezing the runtime so the HTTP server never becomes ready. ### Op bloat 55K ops across only 4,154 sequence numbers (~13 ops per seq on average). Many zeroed-out `MergeJobCrdt` entries appear to be tombstones never cleaned up. Some individual ops are up to 523KB. Average op size is 4.6KB. ## Proposed Fix ### 1. Snapshot (checkpoint) After replaying all ops, serialize the materialized CRDT state to a checkpoint blob (e.g. a `crdt_snapshot` table or a separate file). On next startup, load the snapshot and only replay ops with `rowid > snapshot_rowid`. At snapshot time, back up the database file so corruption is recoverable. ### 2. Op pruning / compaction Delete ops that are superseded by the snapshot. Tombstoned/deleted items with all-zero fields contribute nothing to materialized state and can be dropped from the log once snapshotted. ### 3. Immediate fix: spawn_blocking Move the replay loop to `tokio::task::spawn_blocking` so the HTTP server and liveness ticks are not starved during replay. This does not reduce replay time but prevents the runtime freeze. ## Acceptance Criteria - Startup loads a snapshot when available and only replays ops newer than the snapshot sequence - A snapshot is written after full CRDT replay completes (or periodically in background) - DB backup is created at each snapshot time - Startup time for 55K ops drops from 27+ minutes to under 30 seconds - Dead/zeroed CRDT ops are pruned during compaction - CRDT sync protocol continues to work correctly across nodes after compaction - The replay loop runs in spawn_blocking so it does not freeze the tokio runtime