constitution · epochs · watch · epoch 3

comparison

c_a3b1b1967afc (tommy-mor) vs c_16438843de8f (tommy-mor)

download prompt · raw event · cmp_6c22eaa1ad883a

council reasoning

~anthropic/claude-sonnet-latest · winner B · 8:2 · permalink

Side B adds concrete, functional infrastructure: a working EventLog implementation (append/load with JSONL), a ViewStore for counting views with debounced disk flush, plus deployment artifacts (Dockerfile, fly.toml) and a deps.edn for test tooling. Side A merely deletes a planning markdown document with no accompanying implementation, providing no functional or lasting code value.

~x-ai/grok-latest · winner B · 8:1 · permalink

B adds lasting runtime and ops foundations: EventLog (JSONL append/load), ViewStore persistence, plus Dockerfile/fly.toml deployment. A only deletes PLAN.md with no implementation, tests, or behavioral change—cleanup after planning, not durable project value.

openai/gpt-chat-latest · winner B · 10:1 · permalink

Side B adds functional project infrastructure and code: a JSONL-backed `EventLog` with append/load support, a persistent `ViewStore` that asynchronously flushes counts to disk, plus Docker and Fly.io deployment configuration. Side A only deletes `PLAN.md`, removing design documentation without changing runtime behavior or implementing any of the planned architecture.

sides

A — c_a3b1b1967afc (tommy-mor)

message

[461bfaa7] finished plan

diff preview

diff --git a/PLAN.md b/PLAN.md
deleted file mode 100644
index bc7b3beecea4d63d723432b479a003967f910e45..0000000000000000000000000000000000000000
--- a/PLAN.md
+++ /dev/null
@@ -1,363 +0,0 @@
-# sorter2 storage & memory plan
-
-## Goal
-
-Fit a **decent chunk of Reddit** into a **256MB** Fly VM while keeping the product simple: one Rust binary, no external database service.
-
-**RAM should be bounded by query shape**, not dataset size — ideally one rank-centrality graph in memory at a time, plus runtime overhead.
-
-**`events.jsonl` remains the main database.** Everything on disk elsewhere is a **rebuildable projection**.
-
----
-
-## Architecture (target)
-
-```
-                    ┌─────────────────────────────────┐
-                    │  events.jsonl (source of truth) │
-                    └───────────────┬─────────────────┘
-                                    │
-                         append on every mutation
-                                    │
-                                    ▼
-         ┌──────────────────────────────────────────────┐
-         │  apply event → durable projection (on disk)     │
-         │  (same semantics as today's in-memory reducer)  │
-         └──────────────────────────────────────────────┘
-                    │                    │
-                    ▼                    ▼
-         ┌──────────────────┐   ┌──────────────────────────┐
-         │ entity_payloads  │   │ reducer state per scope   │
-         │ (fat Reddit JSON)│   │ nodes, children, edges, … │
-         └──────────────────┘   └──────────────────────────┘
-                    │                    │
-                    └────────┬───────────┘
-                             ▼
-         ┌──────────────────────────────────────────────┐
-         │  RAM per request (or small LRU cache)         │
-         │  • one scope's GroupState for rank-centrality │
-         │  • children + EntityData (small views)        │
-         │  • RC scratch allocations                     │
-         │  → compute → render → drop / evict            │
-         └──────────────────────────────────────────────┘
-```
-
-This is **event sourcing / CQRS**:
-
-| Layer | Role |
-|-------|------|
-| **JSONL** | Canonical write log; audit; disaster recovery |
-| **Durable (RocksDB)** | Materialized read model + payload store; rebuildable from JSONL |
-| **RAM** | One (or few) hot scopes for ranking and render |
-
-Durable is **not** a second source of truth. If projection and log diverge, **stream JSONL and rebuild durable**.
-
-The plan is sound, but only if the projection layer is treated as a crash-recoverable index:
-
-- JSONL append must mean "the bytes are recoverable after process or VM crash" (`flush` alone is not enough; use `sync_data` / fsync-equivalent on the append path or make an explicit weaker durability tradeoff).
-- Durable projection writes are allowed to lag JSONL, but startup must catch up from the last applied event before serving reads.
-- Event application must be exactly-once for the durable projection. Votes add edge weights, so accidentally replaying the same event twice changes rankings.
-- Projection schema should stay boring and explicit: persist nodes, children, edge weights, voted pairs, and capped recent votes directly. Avoid clever nested abstractions if they make rebuilds, migrations, or audits harder.
-
----
-
-## What we have today (baseline)
-
-| Piece | Status |
-|-------|--------|
-| `events.jsonl` append-only log | ✓ source of truth |
-| `EventLog::replay` streaming one event at a time | ✓ no full `Vec<Event>` at startup |
-| `entity_store` / `entity_db` (RocksDB via `durable`) | ✓ fat payloads off-heap |
-| `EntityData` in `GlobalTree` | ✓ small derived views in RAM |
-| Full `GlobalTree` replayed at boot | ✗ all nodes, all scopes' `GroupState` in RAM |
-| Rank-centrality | ✓ already scoped per parent; reads in-memory `GroupState` |
-
-**Measured RSS (release, ~395 imports, 2.8MB JSONL):**
-
-| Scenario | RSS |
-|----------|-----|
-| Empty data dir | ~14 MB (+ RocksDB baseline) |
-| After boot with data | ~21–22 MB |
-| Pre-offload (in-tree payloads + vec replay) | ~25–34 MB |
-
-Payload offload + streaming replay helped startup peak, but **the full in-memory reducer** is still the scaling ceiling.
-
----
-
-## What lives where (target)
-
-### JSONL (`events.jsonl`)
-
-All mutations, append-only:
-
-- `VoteRecorded` — scope, pair, ratios
-- `EntityImported` — id, full upstream payload
-- `NodeEnsured` — register path
-- (legacy / other event types as present in log)
-
-### Durable / RocksDB (`{data_dir}/…`)
-
-Single embedded DB directory. Collections (names tentative):
-
-| Collection | Contents | Notes |
-|------------|----------|-------|
-| `entity_payloads` | `ItemId → JSON string` | **Done.** Fat Reddit API blobs |
-| `nodes` | `ItemId → { data: EntityData, children: … }` | Small; no raw payload |
-| `scopes/{parent}/…` | `GroupState` materialization | edges, voted_pairs, item_to_idx, recent_votes (capped) |
-
-Nested layout can follow durable's `Map → Map → Vec` patterns (see `durable/docs/001.md` Sorter sketch).
-
-### RAM
-
-| Resident | When |
-|----------|------|
-| Tokio, Axum, reqwest, RocksDB block cache (tuned) | always |
-| **One scope slice** | per request (or LRU of few scopes, byte-capped) |
-| Rank-centrality temporaries | during render for that scope |
-
-**Not** in RAM at steady state: all subreddits, all vote graphs, all payloads.
-
----
-
-## Write path
-
-Order matters:
-
-1. Append event to `events.jsonl` (must durably succeed first)
-2. Apply event to durable projection (same logic as today's `apply_event`)
-3. Invalidate / update in-memory scope cache if that scope is hot
-
-Journal worker already serializes votes disk → tree; extend to **disk → durable** instead of (eventually) **disk → full GlobalTree**.
-
-```rust
-// conceptual
-append(jsonl, event)?;
-apply_to_durable(event)?;
-scope_cache.invalidate(scope_for(event));
-```
-
-On failure after (1): replay from log repairs projection on next boot or via `replay-index` command.
-
-Because JSONL and RocksDB cannot be committed atomically together, durable must record a projection cursor alongside the projection:
-
-- Prefer a monotonically increasing event sequence number in each JSONL event.
-- Acceptable first version: byte offset + line checksum, as long as truncation and partial trailing lines are handled deliberately.
-- Update projection data and cursor in the same RocksDB `WriteBatch`.
-- On startup, read the cursor, scan only the JSONL tail after that cursor, apply missing events, then serve.
-- If the cursor is missing, corrupt, or points past the log, rebuild durable from JSONL.
-
-This keeps the append-first rule simple: if the process dies after JSONL append but before RocksDB apply, catch-up repairs it; if it dies after RocksDB apply but before cursor update, the batch should not expose a cursor that skips work.
-
----
-
-## Read path
-
-For a page under parent scope `P` (e.g. `reddit.com/r/rust`):
-
-1. **Load scope** from durable (or scope LRU hit)
-   - `EntityData` + children for listing
-   - `GroupState` for ranking and pair selection
-2. **Run rank-centrality** on that `GroupState` (requires RAM — that's fine)
-3. **Render**
-4. **Drop** scope from RAM or return to LRU
-
-Payload fetch (rare): `entity_store.get(id)` only when render needs fields not in `EntityData`.
-
----
-
-## Startup & recovery
-
-### Normal startup
-
-```
-open entity_db (RocksDB)
-open / validate scope indexes in same DB
-read projection cursor
-stream only unapplied JSONL tail into durable
-do NOT replay JSONL into RAM
-serve requests (cold scopes loaded on demand)
-```
-
-### Rebuild projection
-
-```
-stream events.jsonl → apply_event → durable
-(one line at a time; same as EventLog::replay today)
-```
-
-Run when:
-
-- First deploy of projection layer
-- Detected corruption / missing durable dir
-- Manual `replay-index` after restoring JSONL from backup
-
-JSONL is the only file you need to trust for recovery.
-
----
-
-## Scope cache (RAM bound)
-
-**Strict mode:** one scope in RAM at a time — simplest, lowest RAM.
-
-**Practical mode:** LRU cache with **byte budget** (e.g. 64–128MB for scopes on a 256MB VM):
-
-- Evict least-recently-used scope's `GroupState` + child views
-- Reload from durable on next visit
-
-Eviction policy is independent of storage engine.
-
----
-
-## Rank-centrality
-
-No change to the algorithm. It already assumes a whole `GroupState` for one parent scope.
-
-Moving reducer to durable does **not** remove RC memory cost — it removes **holding every scope's graph at once**.
-
-Optional later: materialized score vectors on disk, invalidated on vote. Not required for v1 of this plan.
-
----
-
-## Durable mutations (future API)
-
-Separate **intent** from **apply** for batching and testability:
-
-```rust
-let m = rankings.path().key("rust").key(day).push_end(score);
-db.apply(m)?;  // or batch.apply(&[m1, m2, m3])
-```
-
-Benefits:
-
-- One RocksDB `WriteBatch` / one WAL flush per vote or import batch
-- Serializable ops for tests
-- Aligns with JSONL events at the app layer and storage ops at the durable layer
-
-Keep chained `entry().push()` as sugar over `path().…; apply()`.
-
-Type safety: use type-state path builders if we want compile-time nesting; erased `Vec<Op>` only if we accept runtime errors at apply.
-
----
-
-## Storage engine: RocksDB vs alternatives
-
-**Current choice:** RocksDB via vendored `durable/` workspace crate.
-
-| Engine | Verdict for sorter2 |
-|--------|---------------------|
-| **RocksDB** | Good default for LSM, prefix scans, write-heavy votes + bulk imports. C++ dep, tune block cache for 256MB. |
-| **sled** | Pure Rust appeal; production reliability history gives pause. Not a priority switch. |
-| **fjall** | Pure Rust LSM; evaluate with benchmarks if leaving RocksDB. |
-| **redb** | Lighter pure Rust; fine for payload-only store; less ideal for heavy scattered writes across scopes. |
-| **SQLite** | Wrong shape for nested fractal tree; OK for a single KV table only. |
-
-**Switching engines matters less than:**
-
-1. Batched durable writes
-2. Not materializing full reducer in RAM
-3. Scope-local load/evict
-
-RocksDB stays **narrow**: blob attic + materialized reducer projection. Not a replacement for JSONL.
-
----
-
-## Scaling story (before vs after)
-
-| | In-memory reducer (today) | Target (JSONL + durable projection) |
-|--|---------------------------|-------------------------------------|
-| **RAM grows with** | Total nodes + all scopes + (was) payloads | Hot scope count × scope size + runtime |
-| **Disk grows with** | JSONL (+ entity_db payloads today) | JSONL + full durable projection |
-| **Startup** | O(events) replay into RAM | O(1) open DB |
-| **Fails when** | RSS > VM limit | Scope too large for one RC graph, or disk full |
-| **Recovery** | Replay JSONL | Replay JSONL → rebuild durable |
-
-**Rough RAM per active subreddit (~500 posts, moderate votes):**
-
-| Component | Order of magnitude |
-|-----------|-------------------|
-| `EntityData` × 500 | 0.5–2 MB |
-| `GroupState` | 1–5 MB |
-| RC scratch | 1–5 MB |
-| Runtime + tuned RocksDB | 15–25 MB |
-| **Total one hot scope** | **~20–40 MB** |
-
-Multiple subreddits fit on 256MB with LRU eviction, not all resident at once.
-
----
-
-## Implementation phases
-
-### Phase 0 — Done
-
-- [x] Vend `durable` as workspace crate (`durable/`)
-- [x] `entity_store`: payloads in `{data_dir}/entity_db`
-- [x] Remove `entity_raw` from `NodeState`
-- [x] `EventLog::replay`: stream JSONL, one `Event` at a time
-- [x] `apply_event` in `state.rs` for replay semantics
-
-### Phase 1 — Durable projection (write path)
-
-- [ ] Single `Db` under `{data_dir}/store` (payloads + reducer)
-- [ ] JSONL append path uses explicit durable sync semantics (`sync_data` / fsync-equivalent) or documents any weaker mode
-- [ ] Add projection cursor metadata (event sequence, or byte offset + ch

… preview truncated; 2,899 characters omitted

download full diff A

B — c_16438843de8f (tommy-mor)

message

[4cd0d15d] more seed

diff preview

diff --git a/Dockerfile b/Dockerfile
new file mode 100644
index 0000000000000000000000000000000000000000..9cb07c60cb0da063f747cfbf1b3b876ecb8ba03e
--- /dev/null
+++ b/Dockerfile
@@ -0,0 +1,34 @@
+# time 0.3.47+ requires Rust 1.88 (edition 2024)
+FROM rust:1.88-slim as builder
+
+WORKDIR /build
+
+RUN apt-get update && \
+    apt-get install -y pkg-config libssl-dev && \
+    rm -rf /var/lib/apt/lists/*
+
+# Copy source and build. (Keep it simple to avoid remote build cache oddities.)
+COPY . .
+RUN cargo build --release --package slugsocial-server
+
+FROM debian:bookworm-slim
+
+RUN apt-get update && \
+    apt-get install -y ca-certificates && \
+    rm -rf /var/lib/apt/lists/*
+
+WORKDIR /app
+
+COPY --from=builder /build/target/release/slugsocial-server /app/slugsocial-server
+
+# Create data directory for persistent volume
+RUN mkdir -p /data
+
+ENV SLUG_DATA_DIR=/data
+ENV SLUG_EVENT_LOG=/data/events.jsonl
+ENV PORT=8080
+
+EXPOSE 8080
+
+CMD ["/app/slugsocial-server"]
+
diff --git a/deps.edn b/deps.edn
new file mode 100644
index 0000000000000000000000000000000000000000..0bf892d44f491cb2313e01ae8a942c3097c52948
--- /dev/null
+++ b/deps.edn
@@ -0,0 +1,10 @@
+{:paths ["." "test"]
+ :deps {cheshire/cheshire {:mvn/version "5.13.0"}
+        http-kit/http-kit {:mvn/version "2.8.0"}
+        babashka/fs {:mvn/version "0.5.32"}
+        babashka/process {:mvn/version "0.6.25"}
+        com.blockether/spel {:mvn/version "0.7.11"}}
+ :aliases
+ {:kaocha {:extra-deps {lambdaisland/kaocha {:mvn/version "1.91.1392"}
+                        lambdaisland/kaocha-junit-xml {:mvn/version "1.17.101"}}
+          :main-opts ["-m" "kaocha.runner"]}}}
diff --git a/event_log.rs b/event_log.rs
new file mode 100644
index 0000000000000000000000000000000000000000..eaae0d495e43a45d6590603892265a62cc92906e
--- /dev/null
+++ b/event_log.rs
@@ -0,0 +1,83 @@
+use std::path::{Path, PathBuf};
+
+use tokio::{
+    fs::{self, OpenOptions},
+    io::{AsyncBufReadExt, AsyncWriteExt, BufReader},
+};
+
+use crate::events::Event;
+
+#[derive(Debug, thiserror::Error)]
+pub enum EventLogError {
+    #[error("io error: {0}")]
+    Io(#[from] std::io::Error),
+    #[error("json error: {0}")]
+    Json(#[from] serde_json::Error),
+}
+
+#[derive(Debug, Clone)]
+pub struct EventLog {
+    path: PathBuf,
+}
+
+impl EventLog {
+    pub fn new(path: impl Into<PathBuf>) -> Self {
+        Self { path: path.into() }
+    }
+
+    pub fn path(&self) -> &Path {
+        &self.path
+    }
+
+    pub async fn ensure_parent_dir(&self) -> Result<(), EventLogError> {
+        if let Some(parent) = self.path.parent() {
+            fs::create_dir_all(parent).await?;
+        }
+        Ok(())
+    }
+
+    pub async fn append(&self, event: &Event) -> Result<(), EventLogError> {
+        self.ensure_parent_dir().await?;
+        let mut f: tokio::fs::File = OpenOptions::new()
+            .create(true)
+            .append(true)
+            .open(&self.path)
+            .await?;
+
+        let mut line = serde_json::to_string(event)?;
+        line.push('\n');
+        f.write_all(line.as_bytes()).await?;
+        f.flush().await?;
+        Ok(())
+    }
+
+    /// Load events from JSONL. Corrupt lines are skipped and returned as `(line_no, line)`.
+    pub async fn load_all(&self) -> Result<(Vec<Event>, Vec<(usize, String)>), EventLogError> {
+        if !fs::try_exists(&self.path).await? {
+            return Ok((vec![], vec![]));
+        }
+
+        let f = fs::File::open(&self.path).await?;
+        let mut reader = BufReader::new(f).lines();
+
+        let mut events = Vec::new();
+        let mut bad_lines = Vec::new();
+
+        let mut line_no: usize = 0;
+        while let Some(line) = reader.next_line().await? {
+            line_no += 1;
+            let trimmed = line.trim();
+            if trimmed.is_empty() {
+                continue;
+            }
+            match serde_json::from_str::<Event>(trimmed) {
+                Ok(ev) => events.push(ev),
+                Err(_) => bad_lines.push((line_no, line)),
+            }
+        }
+
+        Ok((events, bad_lines))
+    }
+}
+
+
diff --git a/fly.toml b/fly.toml
new file mode 100644
index 0000000000000000000000000000000000000000..bbb9345e527452db1d87a549213645c195eae5fc
--- /dev/null
+++ b/fly.toml
@@ -0,0 +1,42 @@
+app = "slugsocial"
+primary_region = "iad"
+
+[build]
+  dockerfile = "Dockerfile"
+
+[env]
+  SLUG_DATA_DIR = "/data"
+  SLUG_EVENT_LOG = "/data/events.jsonl"
+  PORT = "8080"
+
+[[services]]
+  internal_port = 8080
+  protocol = "tcp"
+
+  [[services.ports]]
+    port = 80
+    handlers = ["http"]
+    force_https = true
+
+  [[services.ports]]
+    port = 443
+    handlers = ["tls", "http"]
+
+  [services.concurrency]
+    type = "connections"
+    hard_limit = 1000
+    soft_limit = 500
+
+  [[services.http_checks]]
+    interval = "10s"
+    timeout = "2s"
+    grace_period = "5s"
+    method = "GET"
+    path = "/healthz"
+    protocol = "http"
+    tls_skip_verify = false
+
+[[mounts]]
+  source = "slugsocial_data"
+  destination = "/data"
+
diff --git a/views.rs b/views.rs
new file mode 100644
index 0000000000000000000000000000000000000000..d4f0ffc49475f014698b4da0de6f476884430813
--- /dev/null
+++ b/views.rs
@@ -0,0 +1,63 @@
+use std::{
+    collections::HashMap,
+    sync::{Arc, Mutex},
+};
+use tokio::sync::mpsc;
+
+type CountMap = Arc<Mutex<HashMap<String, u64>>>;
+
+#[derive(Clone)]
+pub struct ViewStore {
+    counts: CountMap,
+    flush_tx: mpsc::Sender<()>,
+}
+
+impl ViewStore {
+    pub fn new(json_path: &str) -> Self {
+        // Load existing counts from disk on startup (best-effort)
+        let initial: HashMap<String, u64> = std::fs::read_to_string(json_path)
+            .ok()
+            .and_then(|s| serde_json::from_str(&s).ok())
+            .unwrap_or_default();
+
+        let counts: CountMap = Arc::new(Mutex::new(initial));
+        let (flush_tx, mut flush_rx) = mpsc::channel::<()>(64);
+        let path = json_path.to_string();
+
+        let counts_for_writer = counts.clone();
+        tokio::spawn(async move {
+            while flush_rx.recv().await.is_some() {
+                while flush_rx.try_recv().is_ok() {}
+
+                let snapshot: HashMap<String, u64> = {
+                    counts_for_writer.lock().unwrap().clone()
+                };
+
+                let path = path.clone();
+                let _ = tokio::task::spawn_blocking(move || {
+                    if let Ok(json) = serde_json::to_string(&snapshot) {
+                        let tmp = format!("{path}.tmp");
+                        if std::fs::write(&tmp, &json).is_ok() {
+                            let _ = std::fs::rename(&tmp, &path);
+                        }
+                    }
+                })
+                .await;
+            }
+        });
+
+        Self { counts, flush_tx }
+    }
+
+    pub fn increment(&self, path: String) {
+        {
+            let mut map = self.counts.lock().unwrap();
+            *map.entry(path).or_insert(0) += 1;
+        }
+        let _ = self.flush_tx.try_send(());
+    }
+
+    pub fn get_views(&self, path: &str) -> u64 {
+        self.counts.lock().unwrap().get(path).copied().unwrap_or(0)
+    }
+}

download full diff B

Hardlinks — judgments / attempts / prompt

prompt download

judgments

attempts

Prompt text is loaded only by the download route.