You are a constitutional council ranking individual git commits for ownership allocation. Compare these two commits. Decide which contributed more lasting value to the project. Judge substance, not spectacle: - Prefer correct, lasting design and real bugfixes over churn, formatting, renames, or generated noise. - Prefer clarity and necessity over sheer line count. A small precise change can beat a large diffuse one. - Do not favor a side merely because its patch is longer or noisier. - Weight what the change does for the project, not the contributor's name. Return ONLY a JSON object: {"winner": "A" or "B", "ratio": "N:M", "explanation": "..."} The explanation must cite concrete differences in the patches (1-3 sentences). Side A — contributor: tommy-mor Side A — commit message: [1c914c6e] stage set Side A — unified diff (full patch): diff --git a/plan.md b/plan.md new file mode 100644 index 0000000000000000000000000000000000000000..00d6867a1e0ed144a16a020ea037f685ce646c73 --- /dev/null +++ b/plan.md @@ -0,0 +1,155 @@ +# Plan: `ItemId` + `RouteContext` (identity vs hrefs) + +This document is for **the next agent** to continue the refactor without re-deriving context from chat. It supersedes ad-hoc notes: treat it as the checklist of record until the work lands and this file is deleted or trimmed. + +## Goal + +- **Identity** (what lives in the reducer graph, votes, indexes) becomes a **structural `ItemId` enum** in `slug-types`, not a canonical `String` / `CanonicalItemUrl` newtype. +- **Presentation** (tilde / dash display, breadcrumbs) derives from `ItemId` via explicit methods, not string stripping. +- **Routing** (browser `href`s for public vs room) goes through **`RouteContext`** (started in `server/src/html/routing.rs`) so Maud/handlers do not stitch `/r/…` vs `/~` ad hoc. + +**Non-goals for v1 of the migration:** backward-compatible JSONL or dual-read of old canonical strings in the event log (project has accepted breaking changes). If you reintroduce compat, document it here. + +## Current state (as of this plan) + +- **`CanonicalItemUrl`** (`types/src/paths.rs`): newtype around `String`; `parse` / `parent` / `display_path` / `tilde_tail` / etc. Reducer `ContentState`, `VoteData`, ranking, RPC, search, garden, breadcrumbs all use it or `String` keys derived from it. +- **`ThreadNav`** (`server/src/html/forum/nav.rs`): encodes scope prefixes for threads and garden URLs; **`RouteContext`** now wraps `ThreadNav` (`server/src/html/routing.rs`, re-exported from `server/src/html/mod.rs`) but **most HTML still takes `&ThreadNav` directly** — migration incomplete. +- **URL normalization** lives in `types/src/url_normalize.rs` + `canonicalize_item` / `finalize_external_identity_url` in `paths.rs` (YouTube, sorted query params, room path `room_route_segment` in `paths.rs`). +- **Room HTTP paths** are `/r/{short}{slug}` (fused segment); wire **`room_id`** remains `short/slug` for RPC/events. + +## Target architecture + +### `ItemId` (types) + +Suggested shape (adjust after profiling `Ord` / `Hash` / serde size): + +```text +ItemId::Root — tilde ontology root (today `SLUG_TILDE_ONTOLOGY_ROOT`) +ItemId::Local { segments } — slug.social ~/… path as Vec (lowercase segments, non-empty for non-root) +ItemId::External { url: Url } — normalized `url::Url` (crate `url` already in `slug-types`) +``` + +**API surface (minimum):** + +- `ItemId::parse(&str) -> Option` — single entry from DSL / user input / legacy wire (internally may call `canonicalize_item` + structured split). +- `ItemId::to_wire_url(&self) -> String` — only for **external** boundaries if needed (HTTP fetch, rare assertions); avoid using as the primary key once maps use `ItemId`. +- `parent`, `display_path`, `tilde_tail` / `tilde_http_tail`, `tilde_segments`, `last_segment`, `normalized_storage` — port from `CanonicalItemUrl`. +- **`Ord` + `Hash` + `Eq`** stable for `BTreeSet` / `HashMap` (see `write_actor` scope-rank snapshots). +- **`Serialize` / `Deserialize`** — decide **tagged JSON** for any persisted or API-carried structs (e.g. `VoteData` in tests). If RPC must stay stringy for clients, use a **DTO layer** that converts `ItemId` ↔ wire at the boundary only. + +**Remove:** `CanonicalItemUrl` type and all `path_types::CanonicalItemUrl` / `slug_types::paths::CanonicalItemUrl` exports once call sites are migrated. **`Borrow`** on the old newtype goes away; update `nav!` / any code that assumed map keys borrowed as `str`. + +### `RouteContext` (server HTML) + +- **File:** `server/src/html/routing.rs` — **`RouteContext(ThreadNav)`** with `item_href`, `item_href_raw`, `thread_url`, `garden_root_url`, `room_url`, `From`/`Into` `ThreadNav`. +- **Direction:** new code and refactored Maud should take **`&RouteContext`** (or owned where appropriate) instead of `&ThreadNav` when building links. Long term, **`item_href(&ItemId)`** should not parse strings — it should pattern-match `ItemId` and append tilde tail or `/-/…` external tail using the same rules as today’s `ThreadNav::garden_item_url`. + +### Axum / garden routes + +- **No** single catch-all route (explicit decision): keep the existing router layout in `server/src/lib.rs`. +- Room routes stay **`/r/:room_key/...`** with `room_key` fused; parsing via `slug_types::room_id_from_route_segment` / `room_route_segment` in `paths.rs`. + +## Phased execution (recommended order) + +### Phase 0 — Preconditions (quick) + +1. Read **`AGENTS.md`** (UI contract, durability matrix, `RpcCommand` vs `HtmlUiAction`). +2. Run **`cargo test --workspace`** and **`./scripts/clj-test.sh`** on clean `main` before large diffs; repeat after each phase. + +### Phase 1 — `ItemId` in `slug-types` (no server yet) + +1. Add **`ItemId`** (new file e.g. `types/src/item_id.rs` **or** inline at bottom of `paths.rs` — see **Module cycle** below). +2. Implement **`ItemId::parse`** using existing **`canonicalize_item`** + normalization; port **`CanonicalItemUrl`** methods to **`ItemId`** with tests ported from `paths.rs` `#[cfg(test)] mod tests`. +3. **`GardenItemUrl::from_stored(&ItemId, room_wire)`** (and thread helpers) — build absolute hrefs from structure, not from re-parsing a canonical string. +4. **`TildeHttpPathTail::to_item_id`** (rename from `to_canonical`) / **`tilde_http_path_to_item_id`**. +5. **`TildeOntologyPath::from_stored(&ItemId)`**. +6. Export **`ItemId`** from **`types/src/lib.rs`**; update **`server/src/path_types.rs`** re-exports. +7. **Delete `CanonicalItemUrl`** and fix all **in-crate** references in `types` only until `cargo test` passes for `slug-types`. + +**Module cycle trap:** `item_id.rs` must not `use crate::paths::{...}` if `paths.rs` also imports `ItemId` for `GardenItemUrl` in the same module. **Fix one of:** + +- **A)** Put `ItemId` **inside `paths.rs`** below `canonicalize_item` / helpers (simplest, large file), or +- **B)** Split **`canonicalize_item`** (+ dash host helpers + `finalize_external_identity_url`) into **`types/src/item_wire.rs`**, then `paths.rs` + `item_id.rs` both depend on `item_wire` only (cleaner, more files). + +### Phase 2 — Reducer + ranking (server core) + +1. **`server/src/reducer.rs`**: `ContentState` / `GroupState` / **`VoteData`** — replace **`CanonicalItemUrl`** with **`ItemId`** on all maps, sets, deques, vectors. +2. **`apply_vote`**: normalize `a`/`b` via **`ItemId::parse`** or **`ItemId`**-aware logic (remove string round-trip). +3. **`apply_ingest_to_content`**: **`dsl`** still yields strings for item titles in statements; normalize to **`ItemId`** at ingest boundary via **`ItemId::parse`** once per item. +4. **`server/src/ranking.rs`**, **`server/src/scope_rank.rs`**, **`server/src/api/write_actor.rs`** (including **`BTreeSet`** ordering), **`server/src/api/validate.rs`**, **`server/src/api/helpers.rs`** — propagate **`ItemId`**. +5. **`server/tests/basic.rs`** and any reducer tests constructing **`VoteData`** — use **`ItemId::parse(...).unwrap()`** or helpers. + +### Phase 3 — RPC + search + external resolver + +1. **`server/src/api/rpc.rs`**: rank/pair/matchup/search payloads; today many paths use **`GardenItemUrl::from_storage_str(item.as_str(), …)`** — switch to **`ItemId`** + **`GardenItemUrl::from_stored(&item_id, …)`** (or equivalent). +2. **`server/src/html/search.rs`**: scoring uses item path strings — derive from **`ItemId::display_path`** / **`to_wire_url`** only at the scoring boundary if needed. +3. **`server/src/external_resolver.rs`**: take **`&ItemId`** or **`ItemId::external_url()`** instead of **`&CanonicalItemUrl`**. + +### Phase 4 — HTML / Maud + +1. **`ThreadNav::garden_item_url`**: overload or replace with **`garden_item_href(&self, item: &ItemId)`** (no `CanonicalItemUrl::parse` inside). +2. **`RouteContext`**: extend **`item_href(&ItemId)`**; migrate call sites from **`ThreadNav`** to **`RouteContext`** where only link-building is needed (keep **`ThreadNav`** where scope / auth helpers need the full struct). +3. **`server/src/html/garden.rs`**, **`breadcrumb_path.rs`**, **`forum/*`**, **`editor.rs`**: replace **`CanonicalItemUrl`** with **`ItemId`**; breadcrumbs should walk **`ItemId::parent`** without string `rsplit`. +4. **`types` JSON types** (`RankRow`, etc.): decide whether **`GardenItemUrl`** stays string for JSON or becomes a structured field; keep **one** wire format for the public API. + +### Phase 5 — Cleanup + docs + +1. Remove dead **`canonical_path`** / **`breadcrumb_path`** string logic if fully superseded. +2. Update **`AGENTS.md`** if durability, `POST /ui`, or command surfaces change. +3. Delete or shrink **`plan.md`** when done. + +## File / symbol checklist (non-exhaustive — grep-driven) + +Run periodically: + +```bash +rg "CanonicalItemUrl" -g'*.rs' +rg "path_types::CanonicalItemUrl" -g'*.rs' +rg "tilde_http_path_to_canonical" -g'*.rs' +``` + +**High-touch files (from prior exploration):** + +| Area | Files | +|------|--------| +| Types | `types/src/paths.rs`, `types/src/lib.rs`, `types/src/url_normalize.rs`, (optional) `types/src/item_id.rs`, `types/src/item_wire.rs` | +| Server re-exports | `server/src/path_types.rs`, `server/src/canonical_path.rs` | +| Reducer / ingest | `server/src/reducer.rs`, `server/src/dsl.rs` (parse output types if changed) | +| Ranking | `server/src/ranking.rs`, `server/src/scope_rank.rs` | +| Writer / RPC | `server/src/api/write_actor.rs`, `server/src/api/rpc.rs`, `server/src/api/helpers.rs`, `server/src/api/validate.rs` | +| HTML | `server/src/html/garden.rs`, `server/src/html/breadcrumb_path.rs`, `server/src/html/forum/nav.rs`, `server/src/html/routing.rs`, `server/src/html/search.rs`, `server/src/html/editor.rs`, `server/src/html/forum/ingest.rs`, … | +| Tests | `server/tests/basic.rs`, `server/tests/integration.rs`, `types/src/paths.rs` tests, Clojure under `test/` if URLs/assertions mention canonical shapes | + +## Events / JSONL + +- **`Ingest`** events store **`raw` DSL** only — no change required for item identity inside the event. +- If any future event type stores item ids as strings, migrate to **structured `ItemId` serde** or accept string only at the event boundary with immediate parse into **`ItemId`** on `apply_event`. + +## `nav!` macro (`server/src/paths.rs`) + +- Macros use **`keypath($key)`** with **`.clone()`** — **`ItemId`** must be **`Clone`** (already for enums). Remove any reliance on **`Borrow`** for map keys. + +## Testing gate + +After each phase: + +```bash +cargo test --workspace +./scripts/clj-test.sh +``` + +## Risks / gotchas + +1. **`Ord` on `ItemId`**: must match prior **`CanonicalItemUrl`** / `String` ordering wherever **`BTreeSet`** is used (e.g. deterministic scope-rank snapshots in **`write_actor`**). +2. **External `ItemId`**: **`Url`** equality / hashing — normalization is already centralized in **`url_normalize`**; ensure **`ItemId::parse`** always inserts normalized **`Url`** into **`External`**. +3. **Fake parent URLs** in garden (e.g. **`https://.`** for external root ranking): find all **`parse("https://.")`** style hacks and express as **`ItemId`** or a dedicated sentinel. +4. **Serde**: tests and any RPC clients that snapshot JSON may need expectation updates if **`VoteData`** shape changes. + +## Optional follow-ups (not blocking `ItemId`) + +- More **domain normalizers** in **`url_normalize.rs`** (e.g. `music.youtube.com`, Spotify, etc.). +- **Room wire** vs **HTTP segment** helpers already in **`paths.rs`** (`ROOM_SHORT_ID_LEN`, `room_route_segment`, `room_id_from_route_segment`). + +--- + +**End state criteria:** `rg CanonicalItemUrl` returns nothing; reducer maps use **`ItemId`**; HTML link generation for items goes through **`RouteContext` + `ItemId`**; tests and Kaocha green. diff --git a/server/src/html/mod.rs b/server/src/html/mod.rs index c3dccd03dc6da2a2f6f6fa828657e33a951884b1..668403a35c415e6f091362b28c63ac294057fb6d 100644 --- a/server/src/html/mod.rs +++ b/server/src/html/mod.rs @@ -15,6 +15,7 @@ mod breadcrumb_path; mod editor; mod forum; mod garden; +pub mod routing; mod search; pub mod ui_action; use breadcrumb_path::{ExternalOntologyPath, OntologyPath}; @@ -35,6 +36,7 @@ pub use garden::{ external_garden_index, external_ontology_path, garden_index, ontology_path, room_external_garden_index, room_external_ontology_path, room_garden_index, room_ontology_path, }; +pub use routing::RouteContext; pub use search::{search_page, search_results_fragment}; pub use forum::user_profile_page; pub use ui_action::{parse_html_ui_from_form, HtmlUiAction, HtmlUiParseError, UI_RPC_FIELD}; diff --git a/server/src/html/routing.rs b/server/src/html/routing.rs new file mode 100644 index 0000000000000000000000000000000000000000..13e9ae0cbe32d454e34a7737edf8234d2a558502 --- /dev/null +++ b/server/src/html/routing.rs @@ -0,0 +1,71 @@ +//! Scoped browser paths for HTML. **RouteContext** is the intended single place to build `href`s +//! given public vs room scope (the original blueprint name); today it wraps [`ThreadNav`]. +//! +//! Prefer `RouteContext::item_href` / [`RouteContext::thread_url`] in new Maud over stitching +//! `/r/…` vs `/~` manually. Call sites can migrate incrementally from passing `&ThreadNav`. + +use crate::path_types::CanonicalItemUrl; + +use super::forum::ThreadNav; + +#[derive(Clone)] +pub struct RouteContext(ThreadNav); + +impl RouteContext { + #[inline] + pub fn public() -> Self { + Self(ThreadNav::public()) + } + + #[inline] + pub fn from_room_id(room_id: &str) -> Option { + ThreadNav::from_room_id(room_id).map(Self) + } + + #[inline] + pub fn thread_nav(&self) -> &ThreadNav { + &self.0 + } + + #[inline] + pub fn into_thread_nav(self) -> ThreadNav { + self.0 + } + + /// Relative path for a stored canonical item in this scope’s garden. + pub fn item_href(&self, item: &CanonicalItemUrl) -> String { + self.0.garden_item_url(item.as_str()) + } + + /// Same as [`Self::item_href`] but parses `item` first (raw DSL / user paste). + pub fn item_href_raw(&self, item: &str) -> String { + self.0.garden_item_url(item) + } + + #[inline] + pub fn thread_url(&self, tag: &str) -> String { + self.0.thread_url(tag) + } + + #[inline] + pub fn garden_root_url(&self) -> &str { + self.0.garden_root_url() + } + + #[inline] + pub fn room_url(&self) -> &str { + self.0.room_url() + } +} + +impl From for RouteContext { + fn from(nav: ThreadNav) -> Self { + Self(nav) + } +} + +impl From for ThreadNav { + fn from(ctx: RouteContext) -> Self { + ctx.0 + } +} Side B — contributor: tommy-mor Side B — commit message: [4cd0d15d] more seed Side B — unified diff (full patch): diff --git a/Dockerfile b/Dockerfile new file mode 100644 index 0000000000000000000000000000000000000000..9cb07c60cb0da063f747cfbf1b3b876ecb8ba03e --- /dev/null +++ b/Dockerfile @@ -0,0 +1,34 @@ +# time 0.3.47+ requires Rust 1.88 (edition 2024) +FROM rust:1.88-slim as builder + +WORKDIR /build + +RUN apt-get update && \ + apt-get install -y pkg-config libssl-dev && \ + rm -rf /var/lib/apt/lists/* + +# Copy source and build. (Keep it simple to avoid remote build cache oddities.) +COPY . . +RUN cargo build --release --package slugsocial-server + +FROM debian:bookworm-slim + +RUN apt-get update && \ + apt-get install -y ca-certificates && \ + rm -rf /var/lib/apt/lists/* + +WORKDIR /app + +COPY --from=builder /build/target/release/slugsocial-server /app/slugsocial-server + +# Create data directory for persistent volume +RUN mkdir -p /data + +ENV SLUG_DATA_DIR=/data +ENV SLUG_EVENT_LOG=/data/events.jsonl +ENV PORT=8080 + +EXPOSE 8080 + +CMD ["/app/slugsocial-server"] + diff --git a/deps.edn b/deps.edn new file mode 100644 index 0000000000000000000000000000000000000000..0bf892d44f491cb2313e01ae8a942c3097c52948 --- /dev/null +++ b/deps.edn @@ -0,0 +1,10 @@ +{:paths ["." "test"] + :deps {cheshire/cheshire {:mvn/version "5.13.0"} + http-kit/http-kit {:mvn/version "2.8.0"} + babashka/fs {:mvn/version "0.5.32"} + babashka/process {:mvn/version "0.6.25"} + com.blockether/spel {:mvn/version "0.7.11"}} + :aliases + {:kaocha {:extra-deps {lambdaisland/kaocha {:mvn/version "1.91.1392"} + lambdaisland/kaocha-junit-xml {:mvn/version "1.17.101"}} + :main-opts ["-m" "kaocha.runner"]}}} diff --git a/event_log.rs b/event_log.rs new file mode 100644 index 0000000000000000000000000000000000000000..eaae0d495e43a45d6590603892265a62cc92906e --- /dev/null +++ b/event_log.rs @@ -0,0 +1,83 @@ +use std::path::{Path, PathBuf}; + +use tokio::{ + fs::{self, OpenOptions}, + io::{AsyncBufReadExt, AsyncWriteExt, BufReader}, +}; + +use crate::events::Event; + +#[derive(Debug, thiserror::Error)] +pub enum EventLogError { + #[error("io error: {0}")] + Io(#[from] std::io::Error), + #[error("json error: {0}")] + Json(#[from] serde_json::Error), +} + +#[derive(Debug, Clone)] +pub struct EventLog { + path: PathBuf, +} + +impl EventLog { + pub fn new(path: impl Into) -> Self { + Self { path: path.into() } + } + + pub fn path(&self) -> &Path { + &self.path + } + + pub async fn ensure_parent_dir(&self) -> Result<(), EventLogError> { + if let Some(parent) = self.path.parent() { + fs::create_dir_all(parent).await?; + } + Ok(()) + } + + pub async fn append(&self, event: &Event) -> Result<(), EventLogError> { + self.ensure_parent_dir().await?; + let mut f: tokio::fs::File = OpenOptions::new() + .create(true) + .append(true) + .open(&self.path) + .await?; + + let mut line = serde_json::to_string(event)?; + line.push('\n'); + f.write_all(line.as_bytes()).await?; + f.flush().await?; + Ok(()) + } + + /// Load events from JSONL. Corrupt lines are skipped and returned as `(line_no, line)`. + pub async fn load_all(&self) -> Result<(Vec, Vec<(usize, String)>), EventLogError> { + if !fs::try_exists(&self.path).await? { + return Ok((vec![], vec![])); + } + + let f = fs::File::open(&self.path).await?; + let mut reader = BufReader::new(f).lines(); + + let mut events = Vec::new(); + let mut bad_lines = Vec::new(); + + let mut line_no: usize = 0; + while let Some(line) = reader.next_line().await? { + line_no += 1; + let trimmed = line.trim(); + if trimmed.is_empty() { + continue; + } + match serde_json::from_str::(trimmed) { + Ok(ev) => events.push(ev), + Err(_) => bad_lines.push((line_no, line)), + } + } + + Ok((events, bad_lines)) + } +} + + diff --git a/fly.toml b/fly.toml new file mode 100644 index 0000000000000000000000000000000000000000..bbb9345e527452db1d87a549213645c195eae5fc --- /dev/null +++ b/fly.toml @@ -0,0 +1,42 @@ +app = "slugsocial" +primary_region = "iad" + +[build] + dockerfile = "Dockerfile" + +[env] + SLUG_DATA_DIR = "/data" + SLUG_EVENT_LOG = "/data/events.jsonl" + PORT = "8080" + +[[services]] + internal_port = 8080 + protocol = "tcp" + + [[services.ports]] + port = 80 + handlers = ["http"] + force_https = true + + [[services.ports]] + port = 443 + handlers = ["tls", "http"] + + [services.concurrency] + type = "connections" + hard_limit = 1000 + soft_limit = 500 + + [[services.http_checks]] + interval = "10s" + timeout = "2s" + grace_period = "5s" + method = "GET" + path = "/healthz" + protocol = "http" + tls_skip_verify = false + +[[mounts]] + source = "slugsocial_data" + destination = "/data" + diff --git a/views.rs b/views.rs new file mode 100644 index 0000000000000000000000000000000000000000..d4f0ffc49475f014698b4da0de6f476884430813 --- /dev/null +++ b/views.rs @@ -0,0 +1,63 @@ +use std::{ + collections::HashMap, + sync::{Arc, Mutex}, +}; +use tokio::sync::mpsc; + +type CountMap = Arc>>; + +#[derive(Clone)] +pub struct ViewStore { + counts: CountMap, + flush_tx: mpsc::Sender<()>, +} + +impl ViewStore { + pub fn new(json_path: &str) -> Self { + // Load existing counts from disk on startup (best-effort) + let initial: HashMap = std::fs::read_to_string(json_path) + .ok() + .and_then(|s| serde_json::from_str(&s).ok()) + .unwrap_or_default(); + + let counts: CountMap = Arc::new(Mutex::new(initial)); + let (flush_tx, mut flush_rx) = mpsc::channel::<()>(64); + let path = json_path.to_string(); + + let counts_for_writer = counts.clone(); + tokio::spawn(async move { + while flush_rx.recv().await.is_some() { + while flush_rx.try_recv().is_ok() {} + + let snapshot: HashMap = { + counts_for_writer.lock().unwrap().clone() + }; + + let path = path.clone(); + let _ = tokio::task::spawn_blocking(move || { + if let Ok(json) = serde_json::to_string(&snapshot) { + let tmp = format!("{path}.tmp"); + if std::fs::write(&tmp, &json).is_ok() { + let _ = std::fs::rename(&tmp, &path); + } + } + }) + .await; + } + }); + + Self { counts, flush_tx } + } + + pub fn increment(&self, path: String) { + { + let mut map = self.counts.lock().unwrap(); + *map.entry(path).or_insert(0) += 1; + } + let _ = self.flush_tx.try_send(()); + } + + pub fn get_views(&self, path: &str) -> u64 { + self.counts.lock().unwrap().get(path).copied().unwrap_or(0) + } +}