You are a constitutional council ranking individual git commits for ownership allocation. Compare these two commits. Decide which contributed more lasting value to the project. Judge substance, not spectacle: - Prefer correct, lasting design and real bugfixes over churn, formatting, renames, or generated noise. - Prefer clarity and necessity over sheer line count. A small precise change can beat a large diffuse one. - Do not favor a side merely because its patch is longer or noisier. - Weight what the change does for the project, not the contributor's name. Return ONLY a JSON object: {"winner": "A" or "B", "ratio": "N:M", "explanation": "..."} The explanation must cite concrete differences in the patches (1-3 sentences). Side A — contributor: tommy-mor Side A — commit message: [3b3d5873] item refactor Side A — unified diff (full patch): diff --git a/plan.md b/plan.md deleted file mode 100644 index 00d6867a1e0ed144a16a020ea037f685ce646c73..0000000000000000000000000000000000000000 --- a/plan.md +++ /dev/null @@ -1,155 +0,0 @@ -# Plan: `ItemId` + `RouteContext` (identity vs hrefs) - -This document is for **the next agent** to continue the refactor without re-deriving context from chat. It supersedes ad-hoc notes: treat it as the checklist of record until the work lands and this file is deleted or trimmed. - -## Goal - -- **Identity** (what lives in the reducer graph, votes, indexes) becomes a **structural `ItemId` enum** in `slug-types`, not a canonical `String` / `CanonicalItemUrl` newtype. -- **Presentation** (tilde / dash display, breadcrumbs) derives from `ItemId` via explicit methods, not string stripping. -- **Routing** (browser `href`s for public vs room) goes through **`RouteContext`** (started in `server/src/html/routing.rs`) so Maud/handlers do not stitch `/r/…` vs `/~` ad hoc. - -**Non-goals for v1 of the migration:** backward-compatible JSONL or dual-read of old canonical strings in the event log (project has accepted breaking changes). If you reintroduce compat, document it here. - -## Current state (as of this plan) - -- **`CanonicalItemUrl`** (`types/src/paths.rs`): newtype around `String`; `parse` / `parent` / `display_path` / `tilde_tail` / etc. Reducer `ContentState`, `VoteData`, ranking, RPC, search, garden, breadcrumbs all use it or `String` keys derived from it. -- **`ThreadNav`** (`server/src/html/forum/nav.rs`): encodes scope prefixes for threads and garden URLs; **`RouteContext`** now wraps `ThreadNav` (`server/src/html/routing.rs`, re-exported from `server/src/html/mod.rs`) but **most HTML still takes `&ThreadNav` directly** — migration incomplete. -- **URL normalization** lives in `types/src/url_normalize.rs` + `canonicalize_item` / `finalize_external_identity_url` in `paths.rs` (YouTube, sorted query params, room path `room_route_segment` in `paths.rs`). -- **Room HTTP paths** are `/r/{short}{slug}` (fused segment); wire **`room_id`** remains `short/slug` for RPC/events. - -## Target architecture - -### `ItemId` (types) - -Suggested shape (adjust after profiling `Ord` / `Hash` / serde size): - -```text -ItemId::Root — tilde ontology root (today `SLUG_TILDE_ONTOLOGY_ROOT`) -ItemId::Local { segments } — slug.social ~/… path as Vec (lowercase segments, non-empty for non-root) -ItemId::External { url: Url } — normalized `url::Url` (crate `url` already in `slug-types`) -``` - -**API surface (minimum):** - -- `ItemId::parse(&str) -> Option` — single entry from DSL / user input / legacy wire (internally may call `canonicalize_item` + structured split). -- `ItemId::to_wire_url(&self) -> String` — only for **external** boundaries if needed (HTTP fetch, rare assertions); avoid using as the primary key once maps use `ItemId`. -- `parent`, `display_path`, `tilde_tail` / `tilde_http_tail`, `tilde_segments`, `last_segment`, `normalized_storage` — port from `CanonicalItemUrl`. -- **`Ord` + `Hash` + `Eq`** stable for `BTreeSet` / `HashMap` (see `write_actor` scope-rank snapshots). -- **`Serialize` / `Deserialize`** — decide **tagged JSON** for any persisted or API-carried structs (e.g. `VoteData` in tests). If RPC must stay stringy for clients, use a **DTO layer** that converts `ItemId` ↔ wire at the boundary only. - -**Remove:** `CanonicalItemUrl` type and all `path_types::CanonicalItemUrl` / `slug_types::paths::CanonicalItemUrl` exports once call sites are migrated. **`Borrow`** on the old newtype goes away; update `nav!` / any code that assumed map keys borrowed as `str`. - -### `RouteContext` (server HTML) - -- **File:** `server/src/html/routing.rs` — **`RouteContext(ThreadNav)`** with `item_href`, `item_href_raw`, `thread_url`, `garden_root_url`, `room_url`, `From`/`Into` `ThreadNav`. -- **Direction:** new code and refactored Maud should take **`&RouteContext`** (or owned where appropriate) instead of `&ThreadNav` when building links. Long term, **`item_href(&ItemId)`** should not parse strings — it should pattern-match `ItemId` and append tilde tail or `/-/…` external tail using the same rules as today’s `ThreadNav::garden_item_url`. - -### Axum / garden routes - -- **No** single catch-all route (explicit decision): keep the existing router layout in `server/src/lib.rs`. -- Room routes stay **`/r/:room_key/...`** with `room_key` fused; parsing via `slug_types::room_id_from_route_segment` / `room_route_segment` in `paths.rs`. - -## Phased execution (recommended order) - -### Phase 0 — Preconditions (quick) - -1. Read **`AGENTS.md`** (UI contract, durability matrix, `RpcCommand` vs `HtmlUiAction`). -2. Run **`cargo test --workspace`** and **`./scripts/clj-test.sh`** on clean `main` before large diffs; repeat after each phase. - -### Phase 1 — `ItemId` in `slug-types` (no server yet) - -1. Add **`ItemId`** (new file e.g. `types/src/item_id.rs` **or** inline at bottom of `paths.rs` — see **Module cycle** below). -2. Implement **`ItemId::parse`** using existing **`canonicalize_item`** + normalization; port **`CanonicalItemUrl`** methods to **`ItemId`** with tests ported from `paths.rs` `#[cfg(test)] mod tests`. -3. **`GardenItemUrl::from_stored(&ItemId, room_wire)`** (and thread helpers) — build absolute hrefs from structure, not from re-parsing a canonical string. -4. **`TildeHttpPathTail::to_item_id`** (rename from `to_canonical`) / **`tilde_http_path_to_item_id`**. -5. **`TildeOntologyPath::from_stored(&ItemId)`**. -6. Export **`ItemId`** from **`types/src/lib.rs`**; update **`server/src/path_types.rs`** re-exports. -7. **Delete `CanonicalItemUrl`** and fix all **in-crate** references in `types` only until `cargo test` passes for `slug-types`. - -**Module cycle trap:** `item_id.rs` must not `use crate::paths::{...}` if `paths.rs` also imports `ItemId` for `GardenItemUrl` in the same module. **Fix one of:** - -- **A)** Put `ItemId` **inside `paths.rs`** below `canonicalize_item` / helpers (simplest, large file), or -- **B)** Split **`canonicalize_item`** (+ dash host helpers + `finalize_external_identity_url`) into **`types/src/item_wire.rs`**, then `paths.rs` + `item_id.rs` both depend on `item_wire` only (cleaner, more files). - -### Phase 2 — Reducer + ranking (server core) - -1. **`server/src/reducer.rs`**: `ContentState` / `GroupState` / **`VoteData`** — replace **`CanonicalItemUrl`** with **`ItemId`** on all maps, sets, deques, vectors. -2. **`apply_vote`**: normalize `a`/`b` via **`ItemId::parse`** or **`ItemId`**-aware logic (remove string round-trip). -3. **`apply_ingest_to_content`**: **`dsl`** still yields strings for item titles in statements; normalize to **`ItemId`** at ingest boundary via **`ItemId::parse`** once per item. -4. **`server/src/ranking.rs`**, **`server/src/scope_rank.rs`**, **`server/src/api/write_actor.rs`** (including **`BTreeSet`** ordering), **`server/src/api/validate.rs`**, **`server/src/api/helpers.rs`** — propagate **`ItemId`**. -5. **`server/tests/basic.rs`** and any reducer tests constructing **`VoteData`** — use **`ItemId::parse(...).unwrap()`** or helpers. - -### Phase 3 — RPC + search + external resolver - -1. **`server/src/api/rpc.rs`**: rank/pair/matchup/search payloads; today many paths use **`GardenItemUrl::from_storage_str(item.as_str(), …)`** — switch to **`ItemId`** + **`GardenItemUrl::from_stored(&item_id, …)`** (or equivalent). -2. **`server/src/html/search.rs`**: scoring uses item path strings — derive from **`ItemId::display_path`** / **`to_wire_url`** only at the scoring boundary if needed. -3. **`server/src/external_resolver.rs`**: take **`&ItemId`** or **`ItemId::external_url()`** instead of **`&CanonicalItemUrl`**. - -### Phase 4 — HTML / Maud - -1. **`ThreadNav::garden_item_url`**: overload or replace with **`garden_item_href(&self, item: &ItemId)`** (no `CanonicalItemUrl::parse` inside). -2. **`RouteContext`**: extend **`item_href(&ItemId)`**; migrate call sites from **`ThreadNav`** to **`RouteContext`** where only link-building is needed (keep **`ThreadNav`** where scope / auth helpers need the full struct). -3. **`server/src/html/garden.rs`**, **`breadcrumb_path.rs`**, **`forum/*`**, **`editor.rs`**: replace **`CanonicalItemUrl`** with **`ItemId`**; breadcrumbs should walk **`ItemId::parent`** without string `rsplit`. -4. **`types` JSON types** (`RankRow`, etc.): decide whether **`GardenItemUrl`** stays string for JSON or becomes a structured field; keep **one** wire format for the public API. - -### Phase 5 — Cleanup + docs - -1. Remove dead **`canonical_path`** / **`breadcrumb_path`** string logic if fully superseded. -2. Update **`AGENTS.md`** if durability, `POST /ui`, or command surfaces change. -3. Delete or shrink **`plan.md`** when done. - -## File / symbol checklist (non-exhaustive — grep-driven) - -Run periodically: - -```bash -rg "CanonicalItemUrl" -g'*.rs' -rg "path_types::CanonicalItemUrl" -g'*.rs' -rg "tilde_http_path_to_canonical" -g'*.rs' -``` - -**High-touch files (from prior exploration):** - -| Area | Files | -|------|--------| -| Types | `types/src/paths.rs`, `types/src/lib.rs`, `types/src/url_normalize.rs`, (optional) `types/src/item_id.rs`, `types/src/item_wire.rs` | -| Server re-exports | `server/src/path_types.rs`, `server/src/canonical_path.rs` | -| Reducer / ingest | `server/src/reducer.rs`, `server/src/dsl.rs` (parse output types if changed) | -| Ranking | `server/src/ranking.rs`, `server/src/scope_rank.rs` | -| Writer / RPC | `server/src/api/write_actor.rs`, `server/src/api/rpc.rs`, `server/src/api/helpers.rs`, `server/src/api/validate.rs` | -| HTML | `server/src/html/garden.rs`, `server/src/html/breadcrumb_path.rs`, `server/src/html/forum/nav.rs`, `server/src/html/routing.rs`, `server/src/html/search.rs`, `server/src/html/editor.rs`, `server/src/html/forum/ingest.rs`, … | -| Tests | `server/tests/basic.rs`, `server/tests/integration.rs`, `types/src/paths.rs` tests, Clojure under `test/` if URLs/assertions mention canonical shapes | - -## Events / JSONL - -- **`Ingest`** events store **`raw` DSL** only — no change required for item identity inside the event. -- If any future event type stores item ids as strings, migrate to **structured `ItemId` serde** or accept string only at the event boundary with immediate parse into **`ItemId`** on `apply_event`. - -## `nav!` macro (`server/src/paths.rs`) - -- Macros use **`keypath($key)`** with **`.clone()`** — **`ItemId`** must be **`Clone`** (already for enums). Remove any reliance on **`Borrow`** for map keys. - -## Testing gate - -After each phase: - -```bash -cargo test --workspace -./scripts/clj-test.sh -``` - -## Risks / gotchas - -1. **`Ord` on `ItemId`**: must match prior **`CanonicalItemUrl`** / `String` ordering wherever **`BTreeSet`** is used (e.g. deterministic scope-rank snapshots in **`write_actor`**). -2. **External `ItemId`**: **`Url`** equality / hashing — normalization is already centralized in **`url_normalize`**; ensure **`ItemId::parse`** always inserts normalized **`Url`** into **`External`**. -3. **Fake parent URLs** in garden (e.g. **`https://.`** for external root ranking): find all **`parse("https://.")`** style hacks and express as **`ItemId`** or a dedicated sentinel. -4. **Serde**: tests and any RPC clients that snapshot JSON may need expectation updates if **`VoteData`** shape changes. - -## Optional follow-ups (not blocking `ItemId`) - -- More **domain normalizers** in **`url_normalize.rs`** (e.g. `music.youtube.com`, Spotify, etc.). -- **Room wire** vs **HTTP segment** helpers already in **`paths.rs`** (`ROOM_SHORT_ID_LEN`, `room_route_segment`, `room_id_from_route_segment`). - ---- - -**End state criteria:** `rg CanonicalItemUrl` returns nothing; reducer maps use **`ItemId`**; HTML link generation for items goes through **`RouteContext` + `ItemId`**; tests and Kaocha green. diff --git a/server/src/api/helpers.rs b/server/src/api/helpers.rs index 1b291db83df7364a026f2e147e0a29a70a399371..8cd23a02fa5219d6aa375e766e6b2bd2c7bb7dfb 100644 --- a/server/src/api/helpers.rs +++ b/server/src/api/helpers.rs @@ -4,7 +4,8 @@ use axum::{ Json, }; use sha2::{Digest, Sha256}; -use slug_types::paths::{CanonicalItemUrl, GardenItemUrl}; +use slug_types::paths::GardenItemUrl; +use slug_types::ItemId; use slug_types::*; use std::collections::HashMap; @@ -31,12 +32,12 @@ pub fn now_ms() -> i64 { } /// Resolve DSL/user input to a stored canonical item id. -pub fn resolve_item(item: &str) -> Result { +pub fn resolve_item(item: &str) -> Result { let canonical = canonicalize_item(item); if canonical.is_empty() { return Err(format!("empty item path: `{}`", item)); } - Ok(CanonicalItemUrl(canonical)) + ItemId::parse(&canonical).ok_or_else(|| format!("invalid item path: `{}`", item)) } pub fn parse_parent_specs(parent: Option<&String>) -> Vec { @@ -94,7 +95,7 @@ pub fn paginate_rankings( (out_components, out_unranked) } -pub fn pick_random_distinct_canonical(items: &[CanonicalItemUrl]) -> Option<(CanonicalItemUrl, CanonicalItemUrl)> { +pub fn pick_random_distinct_canonical(items: &[ItemId]) -> Option<(ItemId, ItemId)> { use rand::seq::SliceRandom; if items.len() < 2 { return None; @@ -115,15 +116,15 @@ pub fn pick_random_distinct_canonical(items: &[CanonicalItemUrl]) -> Option<(Can } pub fn is_pair_voted(group: &crate::reducer::GroupState, a: &str, b: &str) -> bool { - let a_key = CanonicalItemUrl(a.to_string()); - let b_key = CanonicalItemUrl(b.to_string()); + let a_key = ItemId::parse(a).unwrap_or_else(|| ItemId::opaque(a.to_string())); + let b_key = ItemId::parse(b).unwrap_or_else(|| ItemId::opaque(b.to_string())); let Some(&a_idx) = group.item_to_idx.get(&a_key) else { return false; }; let Some(&b_idx) = group.item_to_idx.get(&b_key) else { return false; }; let (i, j) = if a_idx < b_idx { (a_idx, b_idx) } else { (b_idx, a_idx) }; group.voted_pairs.contains(&(i, j)) } -pub fn compute_connectivity_stats(group: &crate::reducer::GroupState, pool: &[CanonicalItemUrl]) -> ConnectivityStats { +pub fn compute_connectivity_stats(group: &crate::reducer::GroupState, pool: &[ItemId]) -> ConnectivityStats { let n = pool.len(); let global_idxs: Vec> = pool diff --git a/server/src/api/rpc.rs b/server/src/api/rpc.rs index 0ff2d701dd1e8661f58abd55672cd91280e9491e..079d96eb6f1d717c203b1d4aec09f3916384b919 100644 --- a/server/src/api/rpc.rs +++ b/server/src/api/rpc.rs @@ -17,7 +17,7 @@ use crate::{ dsl, events::{Event, Ingest, ThreadCapability}, identity::{parse_agent, parse_username}, - path_types::CanonicalItemUrl, + path_types::ItemId, ranking::{connected_components_from_voted_pairs, ranked_items_subset}, reducer::{scope_from_room_wire, ReducerState, ScopeId}, state::{AppState, InviteState}, @@ -150,7 +150,7 @@ fn build_rank_response_for_content( if !is_global && !specs.is_empty() { let none_exist = specs.iter().all(|spec| { - let Some(canon) = CanonicalItemUrl::parse(spec) else { return true }; + let Some(canon) = ItemId::parse(spec) else { return true }; !content.items.contains(&canon) && !content.item_children.contains_key(&canon) }); if none_exist { @@ -163,10 +163,10 @@ fn build_rank_response_for_content( let depth = depth.max(1); let rankings = if is_global { - let all_items: Vec = content.items.iter().cloned().collect(); + let all_items: Vec = content.items.iter().cloned().collect(); crate::scope_rank::build_rankings_for_item_set(content, &all_items) } else if specs.is_empty() { - crate::scope_rank::build_children_rankings(content, &CanonicalItemUrl::ontology_root()) + crate::scope_rank::build_children_rankings(content, &ItemId::ontology_root()) } else if depth > 1 { let items = crate::scope_rank::resolve_scope_recursive(content, &specs, depth); crate::scope_rank::build_rankings_for_item_set(content, &items) @@ -359,8 +359,8 @@ async fn rpc_check( let mut simulated = { reduced_arc.read().await.clone() }; simulated.apply_event(event); - let voted_parents: Vec = { - let mut parents: HashSet = HashSet::new(); + let voted_parents: Vec = { + let mut parents: HashSet = HashSet::new(); for s in &v.doc.statements { if let dsl::Stmt::Vote { item1, item2, .. } = s { if let (Ok(a), Ok(b)) = (resolve_item(item1), resolve_item(item2)) { @@ -369,7 +369,7 @@ async fn rpc_check( } } } - let mut out: Vec = parents.into_iter().collect(); + let mut out: Vec = parents.into_iter().collect(); out.sort(); out }; @@ -695,7 +695,7 @@ fn rpc_search(reduced: &ReducerState, q: &str, limit: usize, principal: Option<& async fn rpc_get_pair(state: &AppState, room: String, parent_path: String) -> Result { let scope = scope_from_room_wire(&room); let reduced_arc = state.reduced.clone(); - let pool: Vec = { + let pool: Vec = { let reduced = reduced_arc.read().await; let content = content_for_room(&reduced, &room); let tmp = if parent_path.trim().is_empty() { @@ -716,7 +716,7 @@ async fn rpc_get_pair(state: &AppState, room: String, parent_path: String) -> Re Some("add items via ingest".into()), )); } - let selected: Option<(CanonicalItemUrl, CanonicalItemUrl)> = { + let selected: Option<(ItemId, ItemId)> = { let mut reduced = reduced_arc.write().await; let content = reduced.content.entry(scope.clone()).or_default(); let group = &mut content.ranking_group; @@ -729,16 +729,16 @@ async fn rpc_get_pair(state: &AppState, room: String, parent_path: String) -> Re .filter_map(|it| group.item_to_idx.get(it).copied()) .collect(); let ranked = ranked_items_subset(group, &idxs, 10000, 1e-8); - let ranked_set: HashSet = ranked.iter().map(|r| r.item.clone()).collect(); - let unsorted: Vec = pool + let ranked_set: HashSet = ranked.iter().map(|r| r.item.clone()).collect(); + let unsorted: Vec = pool .iter() .filter(|it| !ranked_set.contains(*it)) .cloned() .collect(); - let mut pick: Option<(CanonicalItemUrl, CanonicalItemUrl)> = None; + let mut pick: Option<(ItemId, ItemId)> = None; if !unsorted.is_empty() { if let Some(left) = unsorted.choose(&mut rng).cloned() { - let mut candidates: Vec = if !ranked.is_empty() { + let mut candidates: Vec = if !ranked.is_empty() { ranked.iter().map(|r| r.item.clone()).collect() } else { pool.clone() @@ -865,7 +865,7 @@ pub async fn handle_rpc_batch( } else { let content = content_for_room(&reduced, &room); let item_str = canonicalize_item(&item_path); - let item = CanonicalItemUrl(item_str.clone()); + let item = ItemId::parse(&item_str).unwrap_or_else(|| ItemId::opaque(item_str.clone())); if !content.items.contains(&item) { line_err( "item not found", @@ -1212,7 +1212,7 @@ pub async fn handle_rpc_batch( } let ranked_total = ranked.len(); - let mut unranked: Vec = content + let mut unranked: Vec = content .items .iter() .filter(|it| !group.item_to_idx.contains_key(*it)) @@ -1263,7 +1263,7 @@ pub async fn handle_rpc_batch( } else { let content = content_for_room(&reduced, &room); let item_str = canonicalize_item(&item_path); - let item = CanonicalItemUrl(item_str.clone()); + let item = ItemId::parse(&item_str).unwrap_or_else(|| ItemId::opaque(item_str.clone())); let limit = limit.unwrap_or(50).clamp(1, 200); if !content.items.contains(&item) { line_err( @@ -1304,7 +1304,7 @@ pub async fn handle_rpc_batch( let content = content_for_room(&reduced, &room); let scope = scope_from_room_wire(&room); let item_str = canonicalize_item(&item_path); - let item = CanonicalItemUrl(item_str.clone()); + let item = ItemId::parse(&item_str).unwrap_or_else(|| ItemId::opaque(item_str.clone())); let entries = content.rank_history.get(&item).cloned().unwrap_or_default(); let history: Vec = entries.iter().map(|e| { let caused_by: Vec = reduced.ingests_by_id.get(&e.post_id) @@ -1361,7 +1361,11 @@ pub async fn handle_rpc_batch( line_err(e, h) } else { let content = content_for_room(&reduced, &room); - let parents: HashSet<&str> = content.item_children.keys().map(|s| s.as_str()).collect(); + let parents: HashSet = content + .item_children + .keys() + .map(|k| k.to_storage_string()) + .collect(); let mut paths: Vec = content .items .iter() @@ -1380,11 +1384,11 @@ pub async fn handle_rpc_batch( let content = content_for_room(&reduced, &room); let out: Vec = content .item_children - .get(&CanonicalItemUrl::ontology_root()) + .get(&ItemId::ontology_root()) .map(|roots| { let mut v: Vec = roots.iter() .map(|path| { - let children = content.item_children.get(path.as_str()).map(|s| s.len()).unwrap_or(0); + let children = content.item_children.get(path).map(|s| s.len()).unwrap_or(0); PathSummary { path: TildeOntologyPath::from_stored(path), children, diff --git a/server/src/api/validate.rs b/server/src/api/validate.rs index 3c7aa80fd485cf247231fe5c5c7e6fea62f22534..3f657a88ef9328a31e3e29c3bc3bd32bec6fd7a1 100644 --- a/server/src/api/validate.rs +++ b/server/src/api/validate.rs @@ -4,7 +4,7 @@ use std::collections::HashSet; use crate::{ canonical_path::canonicalize_tag, dsl, - path_types::CanonicalItemUrl, + path_types::ItemId, reducer::{ReducerState, ScopeId}, }; use slug_types::paths::GardenItemUrl; @@ -32,11 +32,11 @@ pub fn validate_ingest_document( ScopeId::Public => None, _ => reduced.content_for_scope(scope), }; - let item_exists = |key: &CanonicalItemUrl| { + let item_exists = |key: &ItemId| { scoped_content.map(|c| c.items.contains(key)).unwrap_or(false) || public_content.items.contains(key) }; - let body_exists = |key: &CanonicalItemUrl| { + let body_exists = |key: &ItemId| { scoped_content.map(|c| c.item_bodies.contains_key(key)).unwrap_or(false) || public_content.item_bodies.contains_key(key) }; @@ -52,7 +52,7 @@ pub fn validate_ingest_document( }; let ts = super::helpers::now_ms(); - let mut defined_in_doc: HashSet = HashSet::new(); + let mut defined_in_doc: HashSet = HashSet::new(); for s in &doc.statements { match s { diff --git a/server/src/api/write_actor.rs b/server/src/api/write_actor.rs index f9c3b8bd3fbf8fcb9c035e1a1572fef0b08fa8a9..3e96b015def54921962f74b4cd884808ff6c3bcc 100644 --- a/server/src/api/write_actor.rs +++ b/server/src/api/write_actor.rs @@ -10,7 +10,7 @@ use crate::{ events::{AgentBound, Event, GrantAdded, Ingest, PostRedacted, RoomDeleted, UserRegistered}, html::JsBuilder, identity::parse_agent, - path_types::CanonicalItemUrl, + path_types::ItemId, reducer::{scope_from_room_wire, ReducerState, ScopeId}, state::AppState, write_cmd::WriteCmd, @@ -90,7 +90,7 @@ async fn broadcast_web_refresh(state: &AppState, room_key: &str, thread_id: &str } fn compute_scope_rank_changes( - parent: &CanonicalItemUrl, + parent: &ItemId, before: &crate::scope_rank::ChildrenRankings, after: &crate::scope_rank::ChildrenRankings, room_wire: &str, @@ -99,7 +99,7 @@ fn compute_scope_rank_changes( use std::collections::BTreeSet; fn build_positions( rankings: &crate::scope_rank::ChildrenRankings, - ) -> HashMap> { + ) -> HashMap> { let mut map = HashMap::new(); for comp in &rankings.component_rankings { let total = comp.ranked.len(); @@ -116,7 +116,7 @@ fn compute_scope_rank_changes( let before_pos = build_positions(before); let after_pos = build_positions(after); - let all_items: BTreeSet = before_pos + let all_items: BTreeSet = before_pos .keys() .cloned() .chain(after_pos.keys().cloned()) @@ -287,8 +287,8 @@ pub async fn writer_actor(mut rx: mpsc::Receiver, state: AppState) { .map(|d| reduced.agent_bindings.get(d).is_none()) .unwrap_or(false); - let voted_parent_scopes: Vec = { - let mut parents: HashSet = HashSet::new(); + let voted_parent_scopes: Vec = { + let mut parents: HashSet = HashSet::new(); for s in &v.doc.statements { if let dsl::Stmt::Vote { item1, item2, .. } = s { if let (Ok(a), Ok(b)) = (resolve_item(item1), resolve_item(item2)) { @@ -301,12 +301,12 @@ pub async fn writer_actor(mut rx: mpsc::Receiver, state: AppState) { } } } - let mut out: Vec = parents.into_iter().collect(); + let mut out: Vec = parents.into_iter().collect(); out.sort(); out }; - let pre_rankings: HashMap = + let pre_rankings: HashMap = if !voted_parent_scopes.is_empty() { let content = content_for_room(&reduced, &room_key); voted_parent_scopes diff --git a/server/src/external_resolver.rs b/server/src/external_resolver.rs index 87409c4cab5f1930d2cd075f3417fa824fd791c7..6f5c982a627f9f620f0e3be0ba0cb92a3d7d4bb7 100644 --- a/server/src/external_resolver.rs +++ b/server/src/external_resolver.rs @@ -1,6 +1,6 @@ use async_trait::async_trait; -use crate::path_types::CanonicalItemUrl; +use crate::path_types::ItemId; #[async_trait] pub trait ExternalResolver: Send + Sync { @@ -11,7 +11,7 @@ pub trait ExternalResolver: Send + Sync { fn normalize(&self, path: &str) -> String; /// Fetches body when missing; GitHub hook lands here in a follow-up. - async fn fetch_body(&self, canonical_url: &CanonicalItemUrl) -> Result; + async fn fetch_body(&self, canonical_url: &ItemId) -> Result; } /// Placeholder until domain-specific resolvers exist. @@ -27,7 +27,7 @@ impl ExternalResolver for DefaultExternalResolver { path.to_string() } - async fn fetch_body(&self, _canonical_url: &CanonicalItemUrl) -> Result { + async fn fetch_body(&self, _canonical_url: &ItemId) -> Result { Err("external fetch not implemented".to_string()) } } diff --git a/server/src/html/breadcrumb_path.rs b/server/src/html/breadcrumb_path.rs index c8a3937923161a6ff248bd77e87eff0dc6fe9ab0..5743a98753f579c70e10961278469afd7cb9ddcf 100644 --- a/server/src/html/breadcrumb_path.rs +++ b/server/src/html/breadcrumb_path.rs @@ -1,8 +1,8 @@ -use crate::path_types::{tilde_http_path_to_canonical, CanonicalItemUrl}; +use crate::path_types::{tilde_http_path_to_item_id, ItemId}; /// Semantic view of an ontology path for rendering and routing decisions. pub(super) struct OntologyPath { - canonical: CanonicalItemUrl, + canonical: ItemId, /// Breadcrumb segments: for `~/a/b` this is `["a", "b"]` (leading `~` rendered separately). segments: Vec, } @@ -11,11 +11,11 @@ impl OntologyPath { /// Path is the `*path` segment from `/~/*path` (e.g. `topic/a`). Always treat it as under `~/` /// so it canonicalizes to `https://slug.social/~/…`, not the non-tilde site path. pub(super) fn from_input(path: &str) -> Self { - let canonical = tilde_http_path_to_canonical(path); + let canonical = tilde_http_path_to_item_id(path); Self::from_canonical(canonical) } - pub(super) fn from_canonical(canonical: CanonicalItemUrl) -> Self { + pub(super) fn from_canonical(canonical: ItemId) -> Self { // tilde_segments() returns ["~", "a", "b"] but bc_path() renders "~" itself, // so we skip the leading "~" segment here. let segments = canonical @@ -28,7 +28,7 @@ impl OntologyPath { } pub(super) fn root() -> Self { - Self::from_canonical(CanonicalItemUrl::ontology_root()) + Self::from_canonical(ItemId::ontology_root()) } pub(super) fn is_root(&self) -> bool { @@ -56,7 +56,7 @@ impl OntologyPath { /// External `https://host/…` items addressed as `/-/host/…` in the URL bar. pub(super) struct ExternalOntologyPath { - canonical: CanonicalItemUrl, + canonical: ItemId, /// e.g. `["github.com", "org", "repo", "issues"]` segments: Vec, } @@ -71,13 +71,13 @@ impl ExternalOntologyPath { } else { format!("-/{}", p.trim_start_matches('/')) }; - let Some(canonical) = CanonicalItemUrl::parse(&raw) else { - return Self::from_canonical(CanonicalItemUrl("https://.".to_string())); + let Some(canonical) = ItemId::parse(&raw) else { + return Self::from_canonical(ItemId::opaque("https://.".to_string())); }; Self::from_canonical(canonical) } - pub(super) fn from_canonical(canonical: CanonicalItemUrl) -> Self { + pub(super) fn from_canonical(canonical: ItemId) -> Self { let s = canonical.as_str(); let rest = s .strip_prefix("https://") diff --git a/server/src/html/editor.rs b/server/src/html/editor.rs index 80657ac0df9b4bb67e4cdacca296f70267840bd1..c76417acd7a164fda23537cd4e7fb9f1f1dcd895 100644 --- a/server/src/html/editor.rs +++ b/server/src/html/editor.rs @@ -97,7 +97,7 @@ pub async fn editor_check( simulated.apply_event(event); // Collect voted parent scopes. - let voted_parents: Vec = { + let voted_parents: Vec = { let mut parents = std::collections::HashSet::new(); for s in &v.doc.statements { if let crate::dsl::Stmt::Vote { item1, item2, .. } = s { @@ -107,7 +107,7 @@ pub async fn editor_check( } } } - let mut out: Vec = parents.into_iter().collect(); + let mut out: Vec = parents.into_iter().collect(); out.sort(); out }; diff --git a/server/src/html/forum/nav.rs b/server/src/html/forum/nav.rs index 48fe11e46731670874ff8b6b05baa6f09ae0b7e4..7ca5b491f38104ec8c81db43a95c59da00abd8c4 100644 --- a/server/src/html/forum/nav.rs +++ b/server/src/html/forum/nav.rs @@ -52,9 +52,14 @@ impl ThreadNav { } pub(crate) fn garden_item_url(&self, item: &str) -> String { - let Some(c) = crate::path_types::CanonicalItemUrl::parse(item) else { + let Some(c) = crate::path_types::ItemId::parse(item) else { return format!("{}/{}", self.garden_path_prefix, canonicalize_item(item)); }; + self.garden_item_href(&c) + } + + /// Relative href for a structured [`crate::path_types::ItemId`] in this scope’s garden. + pub(crate) fn garden_item_href(&self, c: &crate::path_types::ItemId) -> String { if let Some(tail) = c.tilde_tail().map(str::to_owned) { format!("{}/{}", self.garden_path_prefix, tail) } else if c.as_str().starts_with("http://") || c.as_str().starts_with("https://") { @@ -63,7 +68,7 @@ impl ThreadNav { let ext_prefix = format!("{}-", self.garden_path_prefix.trim_end_matches('~')); format!("{}/{}", ext_prefix, rest) } else { - format!("{}/{}", self.garden_path_prefix, canonicalize_item(item)) + format!("{}/{}", self.garden_path_prefix, canonicalize_item(c.as_str())) } } diff --git a/server/src/html/garden.rs b/server/src/html/garden.rs index e615dd356bcf634232d85610c0a26235ead125fd..2af9ac5bfdee1630883e6f8257883baeafcb5a44 100644 --- a/server/src/html/garden.rs +++ b/server/src/html/garden.rs @@ -10,7 +10,7 @@ use crate::{ api::optional_principal, canonical_path::canonicalize_item, events::ThreadCapability, - path_types::CanonicalItemUrl, + path_types::ItemId, reducer::{ContentState, ReducerState, ScopeId}, ranking::{connected_components_from_voted_pairs, ranked_items_subset}, scope_rank::{build_children_rankings, ChildrenRankings}, @@ -27,7 +27,7 @@ use super::{ /// Display path for an item: `~/…` or `-/…` form. fn item_display_path(item: &str) -> String { - CanonicalItemUrl::parse(item) + ItemId::parse(item) .map(|c| c.display_path()) .unwrap_or_else(|| canonicalize_item(item)) } @@ -159,7 +159,7 @@ pub async fn garden_index( let nav = ThreadNav::public(); let child_rankings = { let reduced = state.reduced.read().await; - build_children_rankings(reduced.public(), &CanonicalItemUrl::ontology_root()) + build_children_rankings(reduced.public(), &ItemId::ontology_root()) }; let page = layout( @@ -236,7 +236,7 @@ pub async fn external_garden_index( ) -> impl IntoResponse { let nav = ThreadNav::public(); let ext_path = ExternalOntologyPath::from_input(""); - let parent = CanonicalItemUrl::parse("https://.").unwrap(); + let parent = ItemId::parse("https://.").unwrap(); let child_rankings = { let reduced = state.reduced.read().await; build_children_rankings(reduced.public(), &parent) @@ -365,7 +365,7 @@ pub async fn room_external_garden_index( return room_not_found_page(&jar, &uri).into_response(); } let ext_path = ExternalOntologyPath::from_input(""); - let parent = CanonicalItemUrl::parse("https://.").unwrap(); + let parent = ItemId::parse("https://.").unwrap(); let child_rankings = build_children_rankings( content_for_garden_view(&reduced, &nav.scope()), &parent, @@ -509,13 +509,13 @@ struct ItemPageViewModel { fn build_sibling_rank( reduced: &crate::reducer::ReducerState, scope: &ScopeId, - item: &CanonicalItemUrl, + item: &ItemId, ) -> Option { let item = item.clone().normalized_storage(); let content = content_for_garden_view(reduced, scope); let group = &content.ranking_group; let parent = item.parent()?.normalized_storage(); - let siblings: Vec = content + let siblings: Vec = content .item_children .get(&parent) .map(|s| s.iter().cloned().collect()) @@ -574,7 +574,7 @@ fn build_rank_history( item: &str, ) -> Vec { let content = content_for_garden_view(reduced, scope); - let item_key = CanonicalItemUrl(item.to_string()); + let item_key = ItemId::parse(item).unwrap_or_else(|| ItemId::opaque(item.to_string())); let entries = match content.rank_history.get(&item_key) { None => return vec![], Some(e) => e, @@ -592,8 +592,8 @@ fn build_rank_history( if a_str == item || b_str == item { Some(crate::reducer::VoteData { ts: e.ts, - a: CanonicalItemUrl(a_str), - b: CanonicalItemUrl(b_str), + a: ItemId::parse(&a_str).unwrap_or_else(|| ItemId::opaque(a_str)), + b: ItemId::parse(&b_str).unwrap_or_else(|| ItemId::opaque(b_str)), ratio_left, ratio_right, body: explanation, principal: reduced.ingests_by_id.get(&e.post_id) @@ -629,8 +629,8 @@ fn build_item_page_view_model( item: &str, ) -> ItemPageViewModel { let content = content_for_garden_view(reduced, scope); - let item_key = CanonicalItemUrl::parse(item) - .unwrap_or_else(|| CanonicalItemUrl::parse("~/").unwrap()) + let item_key = ItemId::parse(item) + .unwrap_or_else(|| ItemId::parse("~/").unwrap()) .normalized_storage(); let item_has_parent = item_key.parent().is_some(); let child_rankings = build_children_rankings(content, &item_key); @@ -925,10 +925,10 @@ mod tests { .map(|r| r.item.as_str()) .collect(); assert_eq!(names, vec!["https://slug.social/~/topic/a", "https://slug.social/~/topic/b"]); - use crate::path_types::CanonicalItemUrl; + use crate::path_types::ItemId; assert!( - model.child_rankings.unranked_items.contains(&CanonicalItemUrl("https://slug.social/~/topic/kid1".to_string())) - || model.child_rankings.unranked_items.contains(&CanonicalItemUrl("https://slug.social/~/topic/kid2".to_string())) + model.child_rankings.unranked_items.contains(&ItemId::parse("https://slug.social/~/topic/kid1").unwrap()) + || model.child_rankings.unranked_items.contains(&ItemId::parse("https://slug.social/~/topic/kid2").unwrap()) ); } @@ -941,8 +941,8 @@ mod tests { "9ab12cd/my-room", "@00000000-0000-0000-0000-000000000000:test:local/test\n~/t1 {a}\n~/t2 {b}\n", ); - use crate::path_types::CanonicalItemUrl; - let root = CanonicalItemUrl::ontology_root(); + use crate::path_types::ItemId; + let root = ItemId::ontology_root(); let model = build_item_page_view_model( &reduced, &ScopeId::Room("9ab12cd/my-room".to_string()), @@ -971,8 +971,8 @@ mod tests { "@00000000-0000-0000-0000-000000000000:test:local/test\n\ ~/a {a}\n~/b {b}\n~/a 2:1 ~/b {because}\n", ); - use crate::path_types::CanonicalItemUrl; - let root = CanonicalItemUrl::ontology_root(); + use crate::path_types::ItemId; + let root = ItemId::ontology_root(); let model = build_item_page_view_model( &reduced, &ScopeId::Room("9ab12cd/my-room".to_string()), diff --git a/server/src/html/routing.rs b/server/src/html/routing.rs index 13e9ae0cbe32d454e34a7737edf8234d2a558502..df545a8195aa97484ef0012cff31d802ac67d9df 100644 --- a/server/src/html/routing.rs +++ b/server/src/html/routing.rs @@ -4,7 +4,7 @@ //! Prefer `RouteContext::item_href` / [`RouteContext::thread_url`] in new Maud over stitching //! `/r/…` vs `/~` manually. Call sites can migrate incrementally from passing `&ThreadNav`. -use crate::path_types::CanonicalItemUrl; +use crate::path_types::ItemId; use super::forum::ThreadNav; @@ -32,9 +32,9 @@ impl RouteContext { self.0 } - /// Relative path for a stored canonical item in this scope’s garden. - pub fn item_href(&self, item: &CanonicalItemUrl) -> String { - self.0.garden_item_url(item.as_str()) + /// Relative path for a stored item in this scope’s garden. + pub fn item_href(&self, item: &ItemId) -> String { + self.0.garden_item_href(item) } /// Same as [`Self::item_href`] but parses `item` first (raw DSL / user paste). diff --git a/server/src/path_types.rs b/server/src/path_types.rs index 361a9d446c9043cbdea5f060db2e8633c2fd9bf9..a6e0b90a1aaead0bc1f7e0f4c5296506edf3f7ac 100644 --- a/server/src/path_types.rs +++ b/server/src/path_types.rs @@ -1,5 +1,6 @@ -//! Re-exports — implementations live in `slug-types` (`paths` module). +//! Re-exports — implementations live in `slug-types` (`paths` / `item_id` modules). +pub use slug_types::ItemId; pub use slug_types::paths::{ - tilde_http_path_to_canonical, CanonicalItemUrl, RelativePath, TildeHttpPathTail, TildePath, + tilde_http_path_to_item_id, RelativePath, TildeHttpPathTail, TildePath, }; diff --git a/server/src/ranking.rs b/server/src/ranking.rs index b24e80ddd5f34b4e454bb81b541e8f54bd97a541..3710c9f64437f5bef3b2121905b6f3bcb7611047 100644 --- a/server/src/ranking.rs +++ b/server/src/ranking.rs @@ -1,11 +1,11 @@ use std::collections::HashMap; -use crate::path_types::CanonicalItemUrl; +use crate::path_types::ItemId; use crate::reducer::GroupState; #[derive(Debug, Clone)] pub struct RankedItem { - pub item: CanonicalItemUrl, + pub item: ItemId, pub score: f64, } @@ -239,7 +239,7 @@ pub fn group_summary_scores( group: &mut GroupState, max_iters: usize, tol: f64, -) -> HashMap { +) -> HashMap { ranked_items(group, max_iters, tol) .into_iter() .map(|r| (r.item, r.score)) @@ -256,11 +256,11 @@ mod tests { } fn vote(ts: i64, a: &str, b: &str, l: i32, r: i32) -> VoteData { - use crate::path_types::CanonicalItemUrl; + use crate::path_types::ItemId; VoteData { ts, - a: CanonicalItemUrl(a.to_string()), - b: CanonicalItemUrl(b.to_string()), + a: ItemId::parse(a).unwrap(), + b: ItemId::parse(b).unwrap(), ratio_left: l, ratio_right: r, body: "because".to_string(), diff --git a/server/src/reducer.rs b/server/src/reducer.rs index a8651994f2c062af25b6ad792a014dcc083296f9..cd11a5391de8e36b1c3bb6c9b8067ceba0878e31 100644 --- a/server/src/reducer.rs +++ b/server/src/reducer.rs @@ -5,7 +5,7 @@ use serde::{Deserialize, Serialize}; use crate::canonical_path::canonicalize_tag; use crate::dsl; use crate::events::{Event, Ingest, ThreadCapability}; -use crate::path_types::CanonicalItemUrl; +use crate::path_types::ItemId; #[derive(Debug, Clone, Hash, PartialEq, Eq, PartialOrd, Ord)] pub enum ScopeId { @@ -27,8 +27,8 @@ pub fn scope_from_room_wire(room: &str) -> ScopeId { #[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Eq)] pub struct VoteData { pub ts: i64, - pub a: CanonicalItemUrl, - pub b: CanonicalItemUrl, + pub a: ItemId, + pub b: ItemId, pub ratio_left: i32, pub ratio_right: i32, pub body: String, @@ -40,8 +40,8 @@ pub struct VoteData { #[derive(Debug, Clone)] pub struct GroupState { - pub item_to_idx: HashMap, - pub idx_to_item: Vec, + pub item_to_idx: HashMap, + pub idx_to_item: Vec, /// Aggregated directed edge weights: (src_idx, dst_idx) -> weight. pub edges: HashMap<(usize, usize), f64>, @@ -72,7 +72,7 @@ impl GroupState { } } - fn ensure_item(&mut self, item: &CanonicalItemUrl) -> usize { + fn ensure_item(&mut self, item: &ItemId) -> usize { if let Some(&idx) = self.item_to_idx.get(item) { return idx; } @@ -85,11 +85,11 @@ impl GroupState { /// Public test helper: insert an item into the group without a vote (for unit tests). pub fn ensure_item_pub(&mut self, item: &str) -> usize { - if let Some(canon) = CanonicalItemUrl::parse(item) { + if let Some(canon) = ItemId::parse(item) { self.ensure_item(&canon) } else { // Fallback: treat as raw canonical string - let canon = CanonicalItemUrl(item.to_string()); + let canon = ItemId::opaque(item.to_string()); self.ensure_item(&canon) } } @@ -103,8 +103,8 @@ impl GroupState { } pub fn apply_vote(&mut self, mut vote: VoteData) { - vote.a = CanonicalItemUrl::parse(vote.a.as_str()).unwrap_or(vote.a); - vote.b = CanonicalItemUrl::parse(vote.b.as_str()).unwrap_or(vote.b); + vote.a = ItemId::parse(vote.a.as_str()).unwrap_or_else(|| vote.a.clone()); + vote.b = ItemId::parse(vote.b.as_str()).unwrap_or_else(|| vote.b.clone()); vote.thread_tag = canonicalize_tag(&vote.thread_tag); if vote.ratio_left < 0 { vote.ratio_left = 0; @@ -210,18 +210,18 @@ impl Default for ForumThreadState { #[derive(Debug, Clone, Default)] pub struct ContentState { pub ranking_group: GroupState, - pub items: HashSet, - pub item_bodies: HashMap, + pub items: HashSet, + pub item_bodies: HashMap, /// Parent canonical URL -> direct children. - pub item_children: HashMap>, + pub item_children: HashMap>, /// Per-item vote history (most recent first). - pub item_votes: HashMap>, + pub item_votes: HashMap>, /// Per-item ingest references (most recent first). - pub item_snippets: HashMap>, + pub item_snippets: HashMap>, /// Item path -> threads that mention or vote on this item. - pub item_threads: HashMap>, + pub item_threads: HashMap>, /// Per-item rank history, oldest first. - pub rank_history: HashMap>, + pub rank_history: HashMap>, } #[derive(Debug, Clone)] @@ -351,7 +351,7 @@ impl ReducerState { /// For `a/b/c/d` this creates: `a/b/c→d`, `a/b→a/b/c`, `a→a/b`, `""→a`. /// Stops early when an intermediate is already registered (its ancestors must be too). /// @e2bdefa9-a6fa-4725-b0a2-c0b09d95bb20:claudecode:anthropic/claude-opus-4 - fn add_child_edge(content: &mut ContentState, item: &CanonicalItemUrl) { + fn add_child_edge(content: &mut ContentState, item: &ItemId) { let mut child = item.clone(); loop { let Some(parent) = child.parent() else { break }; @@ -366,16 +366,16 @@ impl ReducerState { } /// Resolve an item path as a first-class canonical path. - fn normalize_item(item: &str) -> Option { - CanonicalItemUrl::parse(item) + fn normalize_item(item: &str) -> Option { + ItemId::parse(item) } /// 1-indexed rank of `item` within its connected component in the parent scope. /// 0 if the item has no votes connecting it to siblings (unranked). fn scope_rank_of( group: &GroupState, - item: &CanonicalItemUrl, - item_children: &HashMap>, + item: &ItemId, + item_children: &HashMap>, ) -> usize { let scope = match item.parent() { Some(p) => p, @@ -421,7 +421,7 @@ impl ReducerState { /// 1-indexed position of `item` in the component-aware global flat list. /// Components sorted largest-first; items ranked within each component. /// 0 if the item is not in the ranking group. - fn global_rank_of(group: &GroupState, item: &CanonicalItemUrl) -> usize { + fn global_rank_of(group: &GroupState, item: &ItemId) -> usize { if !group.item_to_idx.contains_key(item) { return 0; } @@ -448,7 +448,7 @@ impl ReducerState { let doc = dsl::parse_full(&ing.raw).map_err(|_| ())?; let canonical_thread = canonicalize_tag(&ing.thread_tag); - let voted_items: Vec = doc + let voted_items: Vec = doc .statements .iter() .filter_map(|s| { @@ -467,7 +467,7 @@ impl ReducerState { let principal = ing.principal.clone(); let delegate = ing.delegate.clone(); - let before: HashMap = if !voted_items.is_empty() { + let before: HashMap = if !voted_items.is_empty() { crate::ranking::compute_group_ranking(&mut content.ranking_group, 10000, 1e-8); voted_items .iter() @@ -485,7 +485,7 @@ impl ReducerState { HashMap::new() }; - let mut ingest_items: HashSet = HashSet::new(); + let mut ingest_items: HashSet = HashSet::new(); for stmt in doc.statements { match stmt { diff --git a/server/src/scope_rank.rs b/server/src/scope_rank.rs index d656b2f0a0a3623ca6b234b746beaaca8ae6017d..4fdaae150c74ec5162c2accc06d4cc778ad79914 100644 --- a/server/src/scope_rank.rs +++ b/server/src/scope_rank.rs @@ -3,7 +3,7 @@ use std::collections::{HashMap, HashSet}; -use crate::path_types::CanonicalItemUrl; +use crate::path_types::ItemId; use crate::ranking::{connected_components_from_voted_pairs, ranked_items_subset, RankedItem}; use crate::reducer::ContentState; @@ -17,12 +17,12 @@ pub struct ScopedComponent { pub struct ChildrenRankings { pub component_rankings: Vec, /// Items in scope with no rank (no votes connecting them to others in this scope). - pub unranked_items: Vec, + pub unranked_items: Vec, } /// Resolve one scope spec (literal path) to direct children of that parent. No wildcards. -fn resolve_one_scope(content: &ContentState, spec: &str) -> HashSet { - let Some(parent) = CanonicalItemUrl::parse(spec.trim()) else { +fn resolve_one_scope(content: &ContentState, spec: &str) -> HashSet { + let Some(parent) = ItemId::parse(spec.trim()) else { return HashSet::new(); }; content @@ -36,12 +36,12 @@ fn resolve_one_scope(content: &ContentState, spec: &str) -> HashSet Vec { +pub fn resolve_scope(content: &ContentState, specs: &[String]) -> Vec { let mut set = HashSet::new(); for spec in specs { set.extend(resolve_one_scope(content, spec)); } - let mut out: Vec = set.into_iter().collect(); + let mut out: Vec = set.into_iter().collect(); out.sort(); out } @@ -49,18 +49,18 @@ pub fn resolve_scope(content: &ContentState, specs: &[String]) -> Vec Vec { +pub fn resolve_scope_recursive(content: &ContentState, specs: &[String], depth: usize) -> Vec { if depth == 0 { return vec![]; } - let mut visited: HashSet = HashSet::new(); - let mut frontier: Vec = specs + let mut visited: HashSet = HashSet::new(); + let mut frontier: Vec = specs .iter() - .filter_map(|s| CanonicalItemUrl::parse(s)) + .filter_map(|s| ItemId::parse(s)) .collect(); for _level in 0..depth { - let mut next_frontier: Vec = Vec::new(); + let mut next_frontier: Vec = Vec::new(); for parent in &frontier { if let Some(children) = content.item_children.get(parent) { for child in children { @@ -76,16 +76,16 @@ pub fn resolve_scope_recursive(content: &ContentState, specs: &[String], depth: frontier = next_frontier; } - let mut out: Vec = visited.into_iter().collect(); + let mut out: Vec = visited.into_iter().collect(); out.sort(); out } /// Build connected-component rankings for an explicit set of item paths. /// Use this when scope comes from multiple parents (resolve_scope). -pub fn build_rankings_for_item_set(content: &ContentState, items_in_scope: &[CanonicalItemUrl]) -> ChildrenRankings { +pub fn build_rankings_for_item_set(content: &ContentState, items_in_scope: &[ItemId]) -> ChildrenRankings { let group = &content.ranking_group; - let mut items_in_scope: Vec = items_in_scope.to_vec(); + let mut items_in_scope: Vec = items_in_scope.to_vec(); items_in_scope.sort(); let scoped_idxs: Vec = items_in_scope @@ -127,7 +127,7 @@ pub fn build_rankings_for_item_set(content: &ContentState, items_in_scope: &[Can }) .collect(); - let mut unranked_items: Vec = isolate_local_idxs + let mut unranked_items: Vec = isolate_local_idxs .into_iter() .filter_map(|li| local_to_global.get(li).copied()) .filter_map(|idx| group.idx_to_item.get(idx).cloned()) @@ -148,9 +148,9 @@ pub fn build_rankings_for_item_set(content: &ContentState, items_in_scope: &[Can /// Build connected-component rankings for direct children of parent_scope. /// Matches the HTML garden view: multiple components, isolates, no-vote items. -pub fn build_children_rankings(content: &ContentState, parent: &CanonicalItemUrl) -> ChildrenRankings { +pub fn build_children_rankings(content: &ContentState, parent: &ItemId) -> ChildrenRankings { let parent = parent.clone().normalized_storage(); - let items: Vec = content + let items: Vec = content .item_children .get(&parent) .map(|s| s.iter().cloned().collect()) @@ -164,10 +164,13 @@ mod tests { use std::collections::{HashMap, HashSet}; fn content_with_children(edges: &[(&str, &[&str])]) -> ContentState { - let mut item_children: HashMap> = HashMap::new(); + let mut item_children: HashMap> = HashMap::new(); for (parent, children) in edges { - let parent = CanonicalItemUrl((*parent).to_string()); - let set: HashSet = children.iter().map(|s| CanonicalItemUrl((*s).to_string())).collect(); + let parent = ItemId::parse(parent).unwrap(); + let set: HashSet = children + .iter() + .map(|s| ItemId::parse(s).unwrap()) + .collect(); item_children.insert(parent, set); } ContentState { @@ -189,8 +192,8 @@ mod tests { ]); let out = resolve_one_scope(&content, "models"); assert_eq!(out.len(), 2); - assert!(out.contains(&CanonicalItemUrl("https://slug.social/models/x".to_string()))); - assert!(out.contains(&CanonicalItemUrl("https://slug.social/models/y".to_string()))); + assert!(out.contains(&ItemId::parse("https://slug.social/models/x").unwrap())); + assert!(out.contains(&ItemId::parse("https://slug.social/models/y").unwrap())); } #[test] @@ -201,9 +204,9 @@ mod tests { ]); let out = resolve_scope(&content, &["a".into(), "b".into()]); assert_eq!(out.len(), 4); - assert!(out.contains(&CanonicalItemUrl("https://slug.social/a/1".to_string()))); - assert!(out.contains(&CanonicalItemUrl("https://slug.social/a/2".to_string()))); - assert!(out.contains(&CanonicalItemUrl("https://slug.social/b/1".to_string()))); - assert!(out.contains(&CanonicalItemUrl("https://slug.social/b/2".to_string()))); + assert!(out.contains(&ItemId::parse("https://slug.social/a/1").unwrap())); + assert!(out.contains(&ItemId::parse("https://slug.social/a/2").unwrap())); + assert!(out.contains(&ItemId::parse("https://slug.social/b/1").unwrap())); + assert!(out.contains(&ItemId::parse("https://slug.social/b/2").unwrap())); } } diff --git a/server/tests/basic.rs b/server/tests/basic.rs index 17db86894f18d6b542d2943f4cedb0572b247171..c8efea2976e20b443d78a5d4c3b244bc4f2fd6da 100644 --- a/server/tests/basic.rs +++ b/server/tests/basic.rs @@ -7,8 +7,15 @@ use slugsocial_server::{ }; +use slugsocial_server::path_types::ItemId; + use tempfile::TempDir; +#[inline] +fn item_id(s: &str) -> ItemId { + ItemId::parse(s).unwrap() +} + fn ingest_event(ts: i64, raw: &str) -> Event { Event::Ingest(Ingest { ts, @@ -47,7 +54,7 @@ fn reducer_external_namespace_ranking() { let g = &content.ranking_group; assert_eq!(g.idx_to_item.len(), 2); - let parent = slugsocial_server::path_types::CanonicalItemUrl::parse("https://github.com/iss").unwrap(); + let parent = slugsocial_server::path_types::ItemId::parse("https://github.com/iss").unwrap(); let children = slugsocial_server::scope_rank::build_children_rankings(content, &parent); assert_eq!(children.component_rankings.len(), 1); let names: Vec<&str> = children.component_rankings[0] @@ -77,9 +84,9 @@ fn reducer_and_ranking_linear_chain() { let mut group = state.public().ranking_group.clone(); let ranked = ranked_items(&mut group, 20000, 1e-9); assert_eq!(ranked.len(), 3); - assert_eq!(ranked[0].item, "https://slug.social/~/t/a"); - assert_eq!(ranked[1].item, "https://slug.social/~/t/b"); - assert_eq!(ranked[2].item, "https://slug.social/~/t/c"); + assert_eq!(ranked[0].item.as_str(), "https://slug.social/~/t/a"); + assert_eq!(ranked[1].item.as_str(), "https://slug.social/~/t/b"); + assert_eq!(ranked[2].item.as_str(), "https://slug.social/~/t/c"); } #[test] @@ -112,15 +119,15 @@ fn reducer_handles_item_and_body_from_ingest() { )); let content = state.public(); - assert!(content.items.contains("https://slug.social/~/t/test-item")); + assert!(content.items.contains(&item_id("https://slug.social/~/t/test-item"))); assert_eq!( - content.item_bodies.get("https://slug.social/~/t/test-item"), + content.item_bodies.get(&item_id("https://slug.social/~/t/test-item")), Some(&"Description here".to_string()) ); assert!(content .item_children - .get("https://slug.social/~/t") - .map(|c| c.contains("https://slug.social/~/t/test-item")) + .get(&item_id("https://slug.social/~/t")) + .map(|c| c.contains(&item_id("https://slug.social/~/t/test-item"))) .unwrap_or(false)); } @@ -136,12 +143,12 @@ fn reducer_indexes_item_threads_and_vote_thread() { state.apply_event(Event::Ingest(ev)); let content = state.public(); - let threads_for_insertion = content.item_threads.get("https://slug.social/~/sorts/insertion").unwrap(); + let threads_for_insertion = content.item_threads.get(&item_id("https://slug.social/~/sorts/insertion")).unwrap(); assert!(threads_for_insertion.contains("sorting-hat")); - let threads_for_mergesort = content.item_threads.get("https://slug.social/~/sorts/mergesort").unwrap(); + let threads_for_mergesort = content.item_threads.get(&item_id("https://slug.social/~/sorts/mergesort")).unwrap(); assert!(threads_for_mergesort.contains("sorting-hat")); - let vote = content.item_votes.get("https://slug.social/~/sorts/insertion").unwrap().front().unwrap(); + let vote = content.item_votes.get(&item_id("https://slug.social/~/sorts/insertion")).unwrap().front().unwrap(); assert_eq!(vote.thread_tag, "sorting-hat"); } @@ -155,8 +162,8 @@ fn reducer_aggregates_multiple_votes() { } let group = &state.public().ranking_group; - let a_idx = group.item_to_idx["https://slug.social/~/t/a"]; - let b_idx = group.item_to_idx["https://slug.social/~/t/b"]; + let a_idx = group.item_to_idx[&item_id("https://slug.social/~/t/a")]; + let b_idx = group.item_to_idx[&item_id("https://slug.social/~/t/b")]; // Should have accumulated edge weights in both directions. assert!(group.edges.contains_key(&(a_idx, b_idx))); @@ -232,7 +239,7 @@ fn ranking_dominant_item_wins() { let mut group = state.public().ranking_group.clone(); let ranked = ranked_items(&mut group, 20000, 1e-9); - assert_eq!(ranked[0].item, "https://slug.social/~/t/champion"); + assert_eq!(ranked[0].item.as_str(), "https://slug.social/~/t/champion"); assert!(ranked[0].score > ranked[1].score); } @@ -379,7 +386,7 @@ async fn full_workflow_reducer_and_ranking() { let ranked = ranked_items(&mut group, 20000, 1e-9); assert_eq!(ranked.len(), 2); - assert_eq!(ranked[0].item, "https://slug.social/~/langs/rust"); // Should win + assert_eq!(ranked[0].item.as_str(), "https://slug.social/~/langs/rust"); // Should win assert!(ranked[0].score > ranked[1].score); } @@ -395,27 +402,27 @@ fn reducer_materializes_ancestor_path_segments() { // The intermediate path "https://slug.social/~/ai-models/anthropic" should appear as a child of "https://slug.social/~/ai-models". let content = state.public(); - let ai_models_children = content.item_children.get("https://slug.social/~/ai-models").expect("ai-models should have children"); + let ai_models_children = content.item_children.get(&item_id("https://slug.social/~/ai-models")).expect("ai-models should have children"); assert!( - ai_models_children.contains("https://slug.social/~/ai-models/anthropic"), + ai_models_children.contains(&item_id("https://slug.social/~/ai-models/anthropic")), "ai-models/anthropic should be a child of ai-models" ); // The leaf items should still be children of "https://slug.social/~/ai-models/anthropic". - let anthropic_children = content.item_children.get("https://slug.social/~/ai-models/anthropic").expect("ai-models/anthropic should have children"); - assert!(anthropic_children.contains("https://slug.social/~/ai-models/anthropic/claude-opus")); - assert!(anthropic_children.contains("https://slug.social/~/ai-models/anthropic/claude-sonnet")); + let anthropic_children = content.item_children.get(&item_id("https://slug.social/~/ai-models/anthropic")).expect("ai-models/anthropic should have children"); + assert!(anthropic_children.contains(&item_id("https://slug.social/~/ai-models/anthropic/claude-opus"))); + assert!(anthropic_children.contains(&item_id("https://slug.social/~/ai-models/anthropic/claude-sonnet"))); // Root should contain "https://slug.social/~/ai-models". - let root_children = content.item_children.get("https://slug.social/~").expect("root should have children"); - assert!(root_children.contains("https://slug.social/~/ai-models")); + let root_children = content.item_children.get(&item_id("https://slug.social/~")).expect("root should have children"); + assert!(root_children.contains(&item_id("https://slug.social/~/ai-models"))); // The phantom intermediates should NOT be in the items set (they weren't explicitly created). - assert!(!content.items.contains("https://slug.social/~/ai-models")); - assert!(!content.items.contains("https://slug.social/~/ai-models/anthropic")); + assert!(!content.items.contains(&item_id("https://slug.social/~/ai-models"))); + assert!(!content.items.contains(&item_id("https://slug.social/~/ai-models/anthropic"))); // But the leaf items should be. - assert!(content.items.contains("https://slug.social/~/ai-models/anthropic/claude-opus")); - assert!(content.items.contains("https://slug.social/~/ai-models/anthropic/claude-sonnet")); + assert!(content.items.contains(&item_id("https://slug.social/~/ai-models/anthropic/claude-opus"))); + assert!(content.items.contains(&item_id("https://slug.social/~/ai-models/anthropic/claude-sonnet"))); } #[test] @@ -469,9 +476,9 @@ fn ranking_repeated_votes_normalized() { // Same winner regardless of how many times voted. assert_eq!(ranked_once[0].item, ranked_many[0].item); - assert_eq!(ranked_once[0].item, "https://slug.social/~/norm/a"); + assert_eq!(ranked_once[0].item.as_str(), "https://slug.social/~/norm/a"); assert_eq!(ranked_once[1].item, ranked_many[1].item); - assert_eq!(ranked_once[1].item, "https://slug.social/~/norm/b"); + assert_eq!(ranked_once[1].item.as_str(), "https://slug.social/~/norm/b"); // Scores should be identical (normalization makes repeated votes idempotent). let eps = 1e-6; @@ -506,8 +513,8 @@ fn reducer_zero_zero_vote_ratio_normalizes_to_one_one() { "~/t/a {a}\n~/t/b {b}\n~/t/a 0:0 ~/t/b {zero}\n", )); let group = &state.public().ranking_group; - let a_idx = group.item_to_idx["https://slug.social/~/t/a"]; - let b_idx = group.item_to_idx["https://slug.social/~/t/b"]; + let a_idx = group.item_to_idx[&item_id("https://slug.social/~/t/a")]; + let b_idx = group.item_to_idx[&item_id("https://slug.social/~/t/b")]; // 0:0 should normalize to 1:1 — both directions should have weight assert!(group.edges.contains_key(&(a_idx, b_idx))); assert!(group.edges.contains_key(&(b_idx, a_idx))); @@ -523,8 +530,8 @@ fn reducer_negative_ratio_clamped_to_zero() { let mut group = GroupState::new(); group.apply_vote(slugsocial_server::reducer::VoteData { ts: 1, - a: slugsocial_server::path_types::CanonicalItemUrl("https://slug.social/~/t/a".to_string()), - b: slugsocial_server::path_types::CanonicalItemUrl("https://slug.social/~/t/b".to_string()), + a: slugsocial_server::path_types::ItemId::parse("https://slug.social/~/t/a").unwrap(), + b: slugsocial_server::path_types::ItemId::parse("https://slug.social/~/t/b").unwrap(), ratio_left: -5, ratio_right: -3, body: "negative".to_string(), @@ -534,8 +541,8 @@ fn reducer_negative_ratio_clamped_to_zero() { }); assert_eq!(group.idx_to_item.len(), 2); // Both edges should exist (negatives clamped to 0, then 0:0 -> 1:1) - let a_idx = group.item_to_idx["https://slug.social/~/t/a"]; - let b_idx = group.item_to_idx["https://slug.social/~/t/b"]; + let a_idx = group.item_to_idx[&item_id("https://slug.social/~/t/a")]; + let b_idx = group.item_to_idx[&item_id("https://slug.social/~/t/b")]; assert!(group.edges.contains_key(&(a_idx, b_idx))); assert!(group.edges.contains_key(&(b_idx, a_idx))); } @@ -553,24 +560,24 @@ fn reducer_deep_path_ancestor_materialization_four_levels() { let content = state.public(); let tilde_scope = content .item_children - .get("https://slug.social/~") + .get(&item_id("https://slug.social/~")) .expect("~/ scope should have children"); - assert!(tilde_scope.contains("https://slug.social/~/a")); + assert!(tilde_scope.contains(&item_id("https://slug.social/~/a"))); - let a_children = content.item_children.get("https://slug.social/~/a").expect("a should have children"); - assert!(a_children.contains("https://slug.social/~/a/b")); + let a_children = content.item_children.get(&item_id("https://slug.social/~/a")).expect("a should have children"); + assert!(a_children.contains(&item_id("https://slug.social/~/a/b"))); - let ab_children = content.item_children.get("https://slug.social/~/a/b").expect("a/b should have children"); - assert!(ab_children.contains("https://slug.social/~/a/b/c")); + let ab_children = content.item_children.get(&item_id("https://slug.social/~/a/b")).expect("a/b should have children"); + assert!(ab_children.contains(&item_id("https://slug.social/~/a/b/c"))); - let abc_children = content.item_children.get("https://slug.social/~/a/b/c").expect("a/b/c should have children"); - assert!(abc_children.contains("https://slug.social/~/a/b/c/d")); + let abc_children = content.item_children.get(&item_id("https://slug.social/~/a/b/c")).expect("a/b/c should have children"); + assert!(abc_children.contains(&item_id("https://slug.social/~/a/b/c/d"))); // Only the leaf should be in items set - assert!(content.items.contains("https://slug.social/~/a/b/c/d")); - assert!(!content.items.contains("https://slug.social/~/a")); - assert!(!content.items.contains("https://slug.social/~/a/b")); - assert!(!content.items.contains("https://slug.social/~/a/b/c")); + assert!(content.items.contains(&item_id("https://slug.social/~/a/b/c/d"))); + assert!(!content.items.contains(&item_id("https://slug.social/~/a"))); + assert!(!content.items.contains(&item_id("https://slug.social/~/a/b"))); + assert!(!content.items.contains(&item_id("https://slug.social/~/a/b/c"))); } // ============================================================================ @@ -610,7 +617,7 @@ fn ranking_convergence_tolerance_triggers_early_exit() { // Very tight tolerance but huge max_iters — should still converge fast let ranked = ranked_items(&mut group, 1_000_000, 1e-15); assert_eq!(ranked.len(), 2); - assert_eq!(ranked[0].item, "https://slug.social/~/t/a"); + assert_eq!(ranked[0].item.as_str(), "https://slug.social/~/t/a"); } // ============================================================================ @@ -624,14 +631,14 @@ fn test_item_body_overwrite() { 1, "~/t/x {first}\n", )); - assert_eq!(state.public().item_bodies.get("https://slug.social/~/t/x"), Some(&"first".to_string())); + assert_eq!(state.public().item_bodies.get(&item_id("https://slug.social/~/t/x")), Some(&"first".to_string())); state.apply_event(ingest_event( 2, "~/t/x {second}\n", )); assert_eq!( - state.public().item_bodies.get("https://slug.social/~/t/x"), + state.public().item_bodies.get(&item_id("https://slug.social/~/t/x")), Some(&"second".to_string()), "last writer should win for item bodies" ); @@ -645,9 +652,9 @@ fn test_empty_body_not_stored() { "~/t/blank { }\n", )); let content = state.public(); - assert!(content.items.contains("https://slug.social/~/t/blank"), "item should exist"); + assert!(content.items.contains(&item_id("https://slug.social/~/t/blank")), "item should exist"); assert!( - !content.item_bodies.contains_key("https://slug.social/~/t/blank"), + !content.item_bodies.contains_key(&item_id("https://slug.social/~/t/blank")), "whitespace-only body should not be stored" ); } @@ -664,7 +671,7 @@ fn test_duplicate_items_across_ingests() { "~/t/dup {second}\n", )); let content = state.public(); - let count = content.items.iter().filter(|i| *i == "https://slug.social/~/t/dup").count(); + let count = content.items.iter().filter(|i| i.as_str() == "https://slug.social/~/t/dup").count(); assert_eq!(count, 1, "items set should deduplicate across ingests"); } @@ -736,7 +743,7 @@ fn test_thread_id_is_used_for_votes_and_indexes() { let content = state.public(); let vote = content .item_votes - .get("https://slug.social/~/t/a") + .get(&item_id("https://slug.social/~/t/a")) .unwrap() .front() .unwrap(); @@ -745,7 +752,7 @@ fn test_thread_id_is_used_for_votes_and_indexes() { assert!(state .public() .item_threads - .get("https://slug.social/~/t/a") + .get(&item_id("https://slug.social/~/t/a")) .is_some_and(|threads| threads.contains("first"))); } @@ -757,11 +764,11 @@ fn test_rank_history_created_for_voted_items() { "~/t/a {a}\n~/t/b {b}\n~/t/a 3:1 ~/t/b {reason}\n", )); assert!( - state.public().rank_history.contains_key("https://slug.social/~/t/a"), + state.public().rank_history.contains_key(&item_id("https://slug.social/~/t/a")), "rank_history should have entry for voted item a" ); assert!( - state.public().rank_history.contains_key("https://slug.social/~/t/b"), + state.public().rank_history.contains_key(&item_id("https://slug.social/~/t/b")), "rank_history should have entry for voted item b" ); } @@ -774,7 +781,7 @@ fn test_rank_history_not_created_for_unvoted_items() { "~/t/c {just a definition}\n", )); assert!( - !state.public().rank_history.contains_key("https://slug.social/~/t/c"), + !state.public().rank_history.contains_key(&item_id("https://slug.social/~/t/c")), "rank_history should NOT have entry for item with no votes" ); } @@ -786,7 +793,7 @@ fn test_rank_history_first_entry_delta_zero() { 1, "~/t/a {a}\n~/t/b {b}\n~/t/a 3:1 ~/t/b {reason}\n", )); - let history_a = state.public().rank_history.get("https://slug.social/~/t/a").unwrap(); + let history_a = state.public().rank_history.get(&item_id("https://slug.social/~/t/a")).unwrap(); assert_eq!(history_a.len(), 1); assert_eq!( history_a[0].scope_rank_delta, 0, diff --git a/types/src/item_id.rs b/types/src/item_id.rs new file mode 100644 index 0000000000000000000000000000000000000000..5fe9fad39c0d58e546c0a0162eba4348d64e0865 --- /dev/null +++ b/types/src/item_id.rs @@ -0,0 +1,225 @@ +//! Structural item identity for the reducer graph and ranking (vs presentation-only strings). + +use std::cmp::Ordering; +use std::fmt; + +use serde::{Deserialize, Deserializer, Serialize, Serializer}; + +use crate::item_wire::{ + canonicalize_item, external_display_dash_prefix, normalize_slug_ontology_storage_url, + SLUG_TILDE_ONTOLOGY_ROOT, +}; + +/// Structural key for items in [`slug_types`] and the server reducer. +/// +/// Wire / JSON uses the same single string as the former canonical item URL (via serde). +#[derive(Debug, Clone, Hash, PartialEq, Eq)] +pub enum ItemId { + /// Tilde ontology root (`~/`); storage [`SLUG_TILDE_ONTOLOGY_ROOT`]. + Root, + /// `https://slug.social/~/…` (non-root; normalized trailing path). + Local(String), + /// Normalized `http(s)://…` storage form, including non-`~/` paths on `slug.social`. + Web(String), + /// Raw key material that did not round-trip through [`Self::parse`] (historical edge case). + Opaque(String), +} + +impl ItemId { + pub fn parse(input: &str) -> Option { + let c = normalize_slug_ontology_storage_url(&canonicalize_item(input)); + if c.is_empty() { + return None; + } + if c == SLUG_TILDE_ONTOLOGY_ROOT { + return Some(Self::Root); + } + if c.starts_with("https://slug.social/~/") && c != SLUG_TILDE_ONTOLOGY_ROOT { + return Some(Self::Local(c)); + } + Some(Self::Web(c)) + } + + /// Same as the old `ensure_item` fallback: use `s` verbatim as the map key. + pub fn opaque(raw: String) -> Self { + Self::Opaque(raw) + } + + pub fn ontology_root() -> Self { + Self::Root + } + + /// Collapses legacy slug ontology root spellings so [`std::collections::HashMap`] keys match the graph. + pub fn normalized_storage(self) -> Self { + Self::parse(self.as_str()).unwrap_or(self) + } + + pub fn as_str(&self) -> &str { + match self { + ItemId::Root => SLUG_TILDE_ONTOLOGY_ROOT, + ItemId::Local(s) | ItemId::Web(s) | ItemId::Opaque(s) => s, + } + } + + pub fn to_storage_string(&self) -> String { + self.as_str().to_string() + } + + pub fn tilde_tail(&self) -> Option<&str> { + match self { + ItemId::Root => Some(""), + ItemId::Local(s) => s.strip_prefix("https://slug.social/~/"), + ItemId::Web(s) | ItemId::Opaque(s) => { + if let Some(tail) = s.strip_prefix("https://slug.social/~/") { + return Some(tail); + } + if s == SLUG_TILDE_ONTOLOGY_ROOT || s == "https://slug.social/~/" { + return Some(""); + } + None + } + } + } + + /// HTTP garden tail after `~/` (empty at ontology root), or `None` if not under tilde ontology. + pub fn tilde_http_tail(&self) -> Option { + self.tilde_tail().map(str::to_owned) + } + + pub fn last_segment(&self) -> &str { + let s = self.as_str(); + s.rsplit('/').find(|x| !x.is_empty()).unwrap_or(s) + } + + pub fn parent(&self) -> Option { + match self { + ItemId::Root => None, + ItemId::Local(s) => { + if s == SLUG_TILDE_ONTOLOGY_ROOT || s == "https://slug.social/~/" { + return None; + } + let last_slash = s.rfind('/')?; + let parent_str = &s[..last_slash]; + if parent_str.is_empty() { + None + } else { + Self::parse(parent_str) + } + } + ItemId::Web(s) | ItemId::Opaque(s) => { + if s == SLUG_TILDE_ONTOLOGY_ROOT || s == "https://slug.social/~/" { + return None; + } + if let Some(rest) = s.strip_prefix("https://slug.social/~/") { + if rest.is_empty() { + return None; + } + let last_slash = s.rfind('/')?; + let parent_str = &s[..last_slash]; + Self::parse(parent_str) + } else if s.starts_with("https://") { + let rest = s.strip_prefix("https://").unwrap(); + Self::parent_http_url("https://", rest) + } else if s.starts_with("http://") { + let rest = s.strip_prefix("http://").unwrap(); + Self::parent_http_url("http://", rest) + } else { + None + } + } + } + } + + fn parent_http_url(scheme: &'static str, rest: &str) -> Option { + let (host, path) = rest.split_once('/').map_or((rest, ""), |(h, p)| (h, p)); + let host = host.trim(); + let path = path.trim_end_matches('/'); + if path.is_empty() { + return None; + } + let parent_path = path.rsplit_once('/').map(|(p, _)| p).unwrap_or(""); + if parent_path.is_empty() { + Self::parse(&format!("{scheme}{}", host)) + } else { + Self::parse(&format!("{scheme}{}/{}", host, parent_path)) + } + } + + /// `-/` representation for external `https://…` items, `~/…` for slug ontology, else unchanged. + pub fn display_path(&self) -> String { + if let Some(tail) = self.tilde_tail() { + if tail.is_empty() { + return "~/".to_string(); + } + return format!("~/{}", tail); + } + let s = self.as_str(); + if let Some(tail) = s.strip_prefix("https://") { + if tail.starts_with("slug.social") { + return s.to_string(); + } + return external_display_dash_prefix(tail); + } + if let Some(tail) = s.strip_prefix("http://") { + if tail.starts_with("slug.social") { + return s.to_string(); + } + return external_display_dash_prefix(tail); + } + s.to_string() + } + + pub fn tilde_segments(&self) -> Vec<&str> { + match self.tilde_tail() { + Some(tail) if !tail.is_empty() => std::iter::once("~") + .chain(tail.split('/').filter(|s| !s.is_empty())) + .collect(), + Some(_) => vec!["~"], + None => vec![], + } + } + + /// Normalized URL string for HTTP fetch boundaries (external identities). + pub fn to_wire_url(&self) -> String { + self.to_storage_string() + } +} + +impl PartialOrd for ItemId { + fn partial_cmp(&self, other: &Self) -> Option { + Some(self.cmp(other)) + } +} + +impl Ord for ItemId { + fn cmp(&self, other: &Self) -> Ordering { + self.as_str().cmp(other.as_str()) + } +} + +impl fmt::Display for ItemId { + fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { + self.as_str().fmt(f) + } +} + +impl Serialize for ItemId { + fn serialize(&self, serializer: S) -> Result + where + S: Serializer, + { + self.as_str().serialize(serializer) + } +} + +impl<'de> Deserialize<'de> for ItemId { + fn deserialize(deserializer: D) -> Result + where + D: Deserializer<'de>, + { + let s = String::deserialize(deserializer)?; + ItemId::parse(&s).ok_or_else(|| { + serde::de::Error::custom(format!("invalid item id: {s:?}")) + }) + } +} diff --git a/types/src/item_wire.rs b/types/src/item_wire.rs new file mode 100644 index 0000000000000000000000000000000000000000..20e3e7df77a74713c146c7e7222a809173723c7d --- /dev/null +++ b/types/src/item_wire.rs @@ -0,0 +1,184 @@ +//! Wire normalization for item identity strings (`canonicalize_item`, path segments). +//! Split from `paths` so [`crate::item_id::ItemId`] can depend on this without import cycles. + +use crate::url_normalize::{host_preserves_dash_path_case, normalize_http_identity_url}; + +/// Canonical absolute URL for the tilde ontology **root** (`~/` in UI). +pub const SLUG_TILDE_ONTOLOGY_ROOT: &str = "https://slug.social/~"; + +/// Collapse legacy or parser variants of the ontology root to [`SLUG_TILDE_ONTOLOGY_ROOT`]. +pub fn normalize_slug_ontology_storage_url(s: &str) -> String { + if s == "https://slug.social/~/" { + SLUG_TILDE_ONTOLOGY_ROOT.to_string() + } else { + s.to_string() + } +} + +fn finalize_external_identity_url(s: String) -> String { + if s.starts_with("https://slug.social/") { + return s; + } + let normalized = normalize_http_identity_url(&s).unwrap_or_else(|| s.clone()); + strip_redundant_root_slash(&normalized).unwrap_or(normalized) +} + +/// `url::Url` serializes bare hosts with a `/` path; we keep host-only items slash-free for stable +/// keys matching the pre-normalizer spellings. +fn strip_redundant_root_slash(s: &str) -> Option { + let u = url::Url::parse(s).ok()?; + if u.path() == "/" && u.query().is_none() && u.fragment().is_none() { + let scheme = u.scheme(); + let host = u.host_str()?; + return Some(match u.port() { + Some(p) => format!("{scheme}://{host}:{p}"), + None => format!("{scheme}://{host}"), + }); + } + None +} + +/// Ontology item reference → canonical absolute URL on the slug host. +pub fn canonicalize_item(input: &str) -> String { + let s = input.trim(); + if s.is_empty() { + return String::new(); + } + + if let Some(rest) = s.strip_prefix("-/") { + let (host, tail) = rest + .split_once('/') + .map_or((rest, ""), |(h, t)| (h, t)); + let host = host.trim().to_lowercase(); + if host.is_empty() { + return String::new(); + } + let preserve_case = host_preserves_dash_path_case(&host); + return if tail.is_empty() { + finalize_external_identity_url(format!("https://{}", host)) + } else { + let path = tail + .trim_start_matches('/') + .trim_end_matches('/') + .split('/') + .filter_map(|seg| { + let t = seg.trim(); + if t.is_empty() { + None + } else if preserve_case { + Some(t.to_string()) + } else { + Some(t.to_lowercase()) + } + }) + .collect::>() + .join("/"); + finalize_external_identity_url(format!("https://{}/{}", host, path)) + }; + } + + if let Some(rest) = s.strip_prefix("https://") { + let (host, tail) = rest.split_once('/').map_or((rest, ""), |(h, t)| (h, t)); + let host = host.trim().to_lowercase(); + return finalize_external_identity_url(if tail.is_empty() { + format!("https://{}", host) + } else { + format!("https://{}/{}", host, tail) + }); + } + if let Some(rest) = s.strip_prefix("http://") { + let (host, tail) = rest.split_once('/').map_or((rest, ""), |(h, t)| (h, t)); + let host = host.trim().to_lowercase(); + return finalize_external_identity_url(if tail.is_empty() { + format!("http://{}", host) + } else { + format!("http://{}/{}", host, tail) + }); + } + + let is_tilde = s.starts_with("~/"); + let rest = s.strip_prefix("~/").or_else(|| s.strip_prefix("/")).unwrap_or(s); + + let tail = rest + .split('/') + .filter_map(|seg| { + let t = seg.trim(); + if t.is_empty() { + None + } else { + Some(t.to_lowercase()) + } + }) + .collect::>() + .join("/"); + + if is_tilde { + if tail.is_empty() { + return SLUG_TILDE_ONTOLOGY_ROOT.to_string(); + } + format!("https://slug.social/~/{}", tail) + } else if tail.is_empty() { + "https://slug.social".to_string() + } else { + format!("https://slug.social/{}", tail) + } +} + +pub fn item_path_segments(input: &str) -> Vec { + let canonical = canonicalize_item(input); + if canonical.is_empty() { + return vec![]; + } + + if let Some(rest) = canonical.strip_prefix("https://") { + let (host, tail) = rest.split_once('/').map_or((rest, ""), |(h, t)| (h, t)); + let mut out = vec![format!("https://{}", host)]; + out.extend(tail.split('/').filter(|s| !s.is_empty()).map(|s| s.to_string())); + return out; + } + if let Some(rest) = canonical.strip_prefix("http://") { + let (host, tail) = rest.split_once('/').map_or((rest, ""), |(h, t)| (h, t)); + let mut out = vec![format!("http://{}", host)]; + out.extend(tail.split('/').filter(|s| !s.is_empty()).map(|s| s.to_string())); + return out; + } + + canonical + .split('/') + .filter(|s| !s.is_empty()) + .map(|s| s.to_string()) + .collect() +} + +pub fn item_parent_path(input: &str) -> Option { + let segs = item_path_segments(input); + if segs.len() <= 1 { + return None; + } + Some(segs[..segs.len() - 1].join("/")) +} + +pub(crate) fn external_display_dash_prefix(host_and_path: &str) -> String { + let (host, path) = host_and_path + .split_once('/') + .map_or((host_and_path, ""), |(h, p)| (h, p)); + let host = host.trim().to_lowercase(); + let path = path + .trim_end_matches('/') + .split('/') + .filter_map(|seg| { + let t = seg.trim(); + if t.is_empty() { + None + } else { + Some(t.to_lowercase()) + } + }) + .collect::>() + .join("/"); + if path.is_empty() { + format!("-/{}", host) + } else { + format!("-/{}", format!("{}/{}", host, path)) + } +} diff --git a/types/src/lib.rs b/types/src/lib.rs index ceaa574cd32b7e7eb397c9a3191fcbc037b8c606..05bcfcfbf48f1dc7a4c05e32055bd12eb5cfd623 100644 --- a/types/src/lib.rs +++ b/types/src/lib.rs @@ -1,14 +1,20 @@ use serde::{Deserialize, Serialize}; pub mod url_normalize; +pub mod item_wire; +pub mod item_id; pub mod paths; pub mod timeago; +pub use item_id::ItemId; +pub use item_wire::{ + canonicalize_item, item_parent_path, item_path_segments, normalize_slug_ontology_storage_url, + SLUG_TILDE_ONTOLOGY_ROOT, +}; pub use paths::{ - canonicalize_item, canonicalize_tag, item_parent_path, item_path_segments, normalize_slug_ontology_storage_url, - CanonicalItemUrl, ForumThreadUrl, GardenItemUrl, RelativePath, SLUG_TILDE_ONTOLOGY_ROOT, + canonicalize_tag, ForumThreadUrl, GardenItemUrl, RelativePath, room_id_from_route_segment, room_route_segment, ROOM_SHORT_ID_LEN, - TildeHttpPathTail, TildeOntologyPath, TildePath, tilde_http_path_to_canonical, + TildeHttpPathTail, TildeOntologyPath, TildePath, tilde_http_path_to_item_id, }; pub use url_normalize::normalize_http_identity_url; diff --git a/types/src/paths.rs b/types/src/paths.rs index 5c60e762c2ca74d79c74237097d5bcc02ba74af2..a1d066d83e7fbcefc27707dd682e54dbcec1cc56 100644 --- a/types/src/paths.rs +++ b/types/src/paths.rs @@ -3,37 +3,23 @@ //! //! ## String kinds (parse in this module only) //! -//! - **[`canonicalize_item`] / [`CanonicalItemUrl`]** — graph storage key and DSL form; tilde +//! - **[`canonicalize_item`] / [`crate::ItemId`]** — graph storage key and DSL form; tilde //! ontology root is always [`SLUG_TILDE_ONTOLOGY_ROOT`] (no `…/~/` trailing slash only). //! - **[`TildeHttpPathTail`]** — capture from `GET /~/*path` or `…/r/{short}{slug}/~/…` (the `*path` segment). //! - **`-/…` wire form** — external items; see [`canonicalize_item`] dash branch. //! - **[`GardenItemUrl`], [`ForumThreadUrl`]** — JSON / browser href surfaces. //! - **[`ROOM_SHORT_ID_LEN`] / [`room_route_segment`]** — `/r/{short}{slug}` vs wire `short/slug`. -use std::borrow::Borrow; use std::fmt; use std::ops::Deref; use serde::{Deserialize, Serialize}; -use crate::url_normalize::{host_preserves_dash_path_case, normalize_http_identity_url}; - -// --------------------------------------------------------------------------- -// Slug tilde ontology (single storage form for `~/`) -// --------------------------------------------------------------------------- - -/// Canonical absolute URL for the tilde ontology **root** (`~/` in UI). Used as the -/// `item_children` parent key for top-level items and must match [`CanonicalItemUrl::ontology_root`]. -pub const SLUG_TILDE_ONTOLOGY_ROOT: &str = "https://slug.social/~"; - -/// Collapse legacy or parser variants of the ontology root to [`SLUG_TILDE_ONTOLOGY_ROOT`]. -pub fn normalize_slug_ontology_storage_url(s: &str) -> String { - if s == "https://slug.social/~/" { - SLUG_TILDE_ONTOLOGY_ROOT.to_string() - } else { - s.to_string() - } -} +use crate::item_id::ItemId; +pub use crate::item_wire::{ + canonicalize_item, item_parent_path, item_path_segments, normalize_slug_ontology_storage_url, + SLUG_TILDE_ONTOLOGY_ROOT, +}; // --------------------------------------------------------------------------- // Private room HTTP path (`/r/{short}{slug}`; wire id remains `short/slug`) @@ -84,310 +70,10 @@ pub fn canonicalize_tag(input: &str) -> String { input.trim().trim_start_matches('#').to_lowercase() } -fn finalize_external_identity_url(s: String) -> String { - if s.starts_with("https://slug.social/") { - return s; - } - let normalized = normalize_http_identity_url(&s).unwrap_or_else(|| s.clone()); - strip_redundant_root_slash(&normalized).unwrap_or(normalized) -} - -/// `url::Url` serializes bare hosts with a `/` path; we keep host-only items slash-free for stable -/// keys matching the pre-normalizer spellings. -fn strip_redundant_root_slash(s: &str) -> Option { - let u = url::Url::parse(s).ok()?; - if u.path() == "/" && u.query().is_none() && u.fragment().is_none() { - let scheme = u.scheme(); - let host = u.host_str()?; - return Some(match u.port() { - Some(p) => format!("{scheme}://{host}:{p}"), - None => format!("{scheme}://{host}"), - }); - } - None -} - -/// Ontology item reference → canonical absolute URL on the slug host. -pub fn canonicalize_item(input: &str) -> String { - let s = input.trim(); - if s.is_empty() { - return String::new(); - } - - // External scope: `-/host/path` is the universal alias for `https://host/path`. - if let Some(rest) = s.strip_prefix("-/") { - let (host, tail) = rest - .split_once('/') - .map_or((rest, ""), |(h, t)| (h, t)); - let host = host.trim().to_lowercase(); - if host.is_empty() { - return String::new(); - } - let preserve_case = host_preserves_dash_path_case(&host); - return if tail.is_empty() { - finalize_external_identity_url(format!("https://{}", host)) - } else { - let path = tail - .trim_start_matches('/') - .trim_end_matches('/') - .split('/') - .filter_map(|seg| { - let t = seg.trim(); - if t.is_empty() { - None - } else if preserve_case { - Some(t.to_string()) - } else { - Some(t.to_lowercase()) - } - }) - .collect::>() - .join("/"); - finalize_external_identity_url(format!("https://{}/{}", host, path)) - }; - } - - if let Some(rest) = s.strip_prefix("https://") { - let (host, tail) = rest.split_once('/').map_or((rest, ""), |(h, t)| (h, t)); - let host = host.trim().to_lowercase(); - return finalize_external_identity_url(if tail.is_empty() { - format!("https://{}", host) - } else { - format!("https://{}/{}", host, tail) - }); - } - if let Some(rest) = s.strip_prefix("http://") { - let (host, tail) = rest.split_once('/').map_or((rest, ""), |(h, t)| (h, t)); - let host = host.trim().to_lowercase(); - return finalize_external_identity_url(if tail.is_empty() { - format!("http://{}", host) - } else { - format!("http://{}/{}", host, tail) - }); - } - - let is_tilde = s.starts_with("~/"); - let rest = s.strip_prefix("~/").or_else(|| s.strip_prefix("/")).unwrap_or(s); - - let tail = rest - .split('/') - .filter_map(|seg| { - let t = seg.trim(); - if t.is_empty() { - None - } else { - Some(t.to_lowercase()) - } - }) - .collect::>() - .join("/"); - - if is_tilde { - if tail.is_empty() { - return SLUG_TILDE_ONTOLOGY_ROOT.to_string(); - } - format!("https://slug.social/~/{}", tail) - } else if tail.is_empty() { - "https://slug.social".to_string() - } else { - format!("https://slug.social/{}", tail) - } -} - -pub fn item_path_segments(input: &str) -> Vec { - let canonical = canonicalize_item(input); - if canonical.is_empty() { - return vec![]; - } - - if let Some(rest) = canonical.strip_prefix("https://") { - let (host, tail) = rest.split_once('/').map_or((rest, ""), |(h, t)| (h, t)); - let mut out = vec![format!("https://{}", host)]; - out.extend(tail.split('/').filter(|s| !s.is_empty()).map(|s| s.to_string())); - return out; - } - if let Some(rest) = canonical.strip_prefix("http://") { - let (host, tail) = rest.split_once('/').map_or((rest, ""), |(h, t)| (h, t)); - let mut out = vec![format!("http://{}", host)]; - out.extend(tail.split('/').filter(|s| !s.is_empty()).map(|s| s.to_string())); - return out; - } - - canonical - .split('/') - .filter(|s| !s.is_empty()) - .map(|s| s.to_string()) - .collect() -} - -pub fn item_parent_path(input: &str) -> Option { - let segs = item_path_segments(input); - if segs.len() <= 1 { - return None; - } - Some(segs[..segs.len() - 1].join("/")) -} - -fn external_display_dash_prefix(host_and_path: &str) -> String { - let (host, path) = host_and_path - .split_once('/') - .map_or((host_and_path, ""), |(h, p)| (h, p)); - let host = host.trim().to_lowercase(); - let path = path - .trim_end_matches('/') - .split('/') - .filter_map(|seg| { - let t = seg.trim(); - if t.is_empty() { - None - } else { - Some(t.to_lowercase()) - } - }) - .collect::>() - .join("/"); - if path.is_empty() { - format!("-/{}", host) - } else { - format!("-/{}", format!("{}/{}", host, path)) - } -} - // --------------------------------------------------------------------------- // Storage + input path newtypes // --------------------------------------------------------------------------- -/// Canonical item identifier as produced by [`canonicalize_item`]. -/// -/// Shared across all scopes; room is not embedded. Usually -/// `https://slug.social/~/…` or an external `http(s)://…` URL item. -#[derive(Debug, Clone, PartialEq, Eq, PartialOrd, Ord, Hash, Serialize, Deserialize)] -pub struct CanonicalItemUrl(pub String); - -impl CanonicalItemUrl { - pub fn parse(input: &str) -> Option { - let c = canonicalize_item(input); - if c.is_empty() { - None - } else { - Some(Self(normalize_slug_ontology_storage_url(&c))) - } - } - - pub fn as_str(&self) -> &str { - &self.0 - } - - /// Collapses legacy slug ontology root spellings so [`HashMap`] keys match the reducer graph. - pub fn normalized_storage(self) -> Self { - Self(normalize_slug_ontology_storage_url(self.as_str())) - } - - pub fn tilde_tail(&self) -> Option<&str> { - if let Some(tail) = self.0.strip_prefix("https://slug.social/~/") { - return Some(tail); - } - if self.0 == SLUG_TILDE_ONTOLOGY_ROOT || self.0 == "https://slug.social/~/" { - return Some(""); - } - None - } - - pub fn last_segment(&self) -> &str { - self.0 - .rsplit('/') - .find(|s| !s.is_empty()) - .unwrap_or(self.0.as_str()) - } - - pub fn ontology_root() -> Self { - Self(SLUG_TILDE_ONTOLOGY_ROOT.to_string()) - } - - pub fn parent(&self) -> Option { - if self.tilde_tail().is_some() { - if self.tilde_tail().map(|t| t.is_empty()).unwrap_or(true) { - return None; - } - let last_slash = self.0.rfind('/')?; - let parent_str = &self.0[..last_slash]; - if parent_str.is_empty() { - None - } else { - Some(Self(parent_str.to_string())) - } - } else if let Some(rest) = self.0.strip_prefix("https://") { - Self::parent_http_url("https://", rest) - } else if let Some(rest) = self.0.strip_prefix("http://") { - Self::parent_http_url("http://", rest) - } else { - None - } - } - - fn parent_http_url(scheme: &'static str, rest: &str) -> Option { - let (host, path) = rest.split_once('/').map_or((rest, ""), |(h, p)| (h, p)); - let host = host.trim(); - let path = path.trim_end_matches('/'); - if path.is_empty() { - return None; - } - let parent_path = path.rsplit_once('/').map(|(p, _)| p).unwrap_or(""); - if parent_path.is_empty() { - Some(Self(format!("{scheme}{}", host))) - } else { - Some(Self(format!("{scheme}{}/{}", host, parent_path))) - } - } - - /// `-/` representation for external `https://…` items, `~/…` for slug ontology, else unchanged. - pub fn display_path(&self) -> String { - if let Some(tail) = self.tilde_tail() { - if tail.is_empty() { - return "~/".to_string(); - } - return format!("~/{}", tail); - } - if let Some(tail) = self.0.strip_prefix("https://") { - if tail.starts_with("slug.social") { - self.0.clone() - } else { - external_display_dash_prefix(tail) - } - } else if let Some(tail) = self.0.strip_prefix("http://") { - if tail.starts_with("slug.social") { - self.0.clone() - } else { - external_display_dash_prefix(tail) - } - } else { - self.0.clone() - } - } - - pub fn tilde_segments(&self) -> Vec<&str> { - match self.tilde_tail() { - Some(tail) if !tail.is_empty() => { - std::iter::once("~") - .chain(tail.split('/').filter(|s| !s.is_empty())) - .collect() - } - Some(_) => vec!["~"], - None => vec![], - } - } - - /// `~/…` list label for ontology items (paths index, CLI). - pub fn tilde_list_label(&self) -> TildeOntologyPath { - TildeOntologyPath::from_stored(self) - } - - /// Absolute href for JSON/RPC and browsers for this stored id in `room`. - pub fn json_href(&self, room_wire: &str) -> GardenItemUrl { - GardenItemUrl::from_stored(self, room_wire) - } -} - /// HTTP route capture: path segment after `~/` in `GET /~/*path` or `…/r/{short}{slug}/~/…` (empty = ontology root). #[derive(Debug, Clone, PartialEq, Eq, Hash)] pub struct TildeHttpPathTail(pub String); @@ -401,13 +87,13 @@ impl TildeHttpPathTail { &self.0 } - pub fn to_canonical(&self) -> CanonicalItemUrl { - tilde_http_path_to_canonical(self.as_str()) + pub fn to_item_id(&self) -> ItemId { + tilde_http_path_to_item_id(self.as_str()) } } -/// Map the router's tilde tail (e.g. `topic/a`, or empty for root) to a [`CanonicalItemUrl`]. -pub fn tilde_http_path_to_canonical(path_segment: &str) -> CanonicalItemUrl { +/// Map the router's tilde tail (e.g. `topic/a`, or empty for root) to an [`ItemId`]. +pub fn tilde_http_path_to_item_id(path_segment: &str) -> ItemId { let p = path_segment.trim_start_matches('/'); let raw = if p.starts_with("http://") || p.starts_with("https://") { p.to_string() @@ -416,37 +102,7 @@ pub fn tilde_http_path_to_canonical(path_segment: &str) -> CanonicalItemUrl { } else { format!("~/{}", p) }; - CanonicalItemUrl::parse(&raw).unwrap_or_else(|| CanonicalItemUrl::ontology_root()) -} - -impl fmt::Display for CanonicalItemUrl { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - self.0.fmt(f) - } -} - -impl Borrow for CanonicalItemUrl { - fn borrow(&self) -> &str { - &self.0 - } -} - -impl PartialEq for CanonicalItemUrl { - fn eq(&self, other: &str) -> bool { - self.0 == other - } -} - -impl PartialEq<&str> for CanonicalItemUrl { - fn eq(&self, other: &&str) -> bool { - self.0 == *other - } -} - -impl PartialEq for CanonicalItemUrl { - fn eq(&self, other: &String) -> bool { - &self.0 == other - } + ItemId::parse(&raw).unwrap_or_else(ItemId::ontology_root) } #[derive(Debug, Clone, PartialEq, Eq, PartialOrd, Ord, Hash, Serialize, Deserialize)] @@ -468,8 +124,8 @@ impl TildePath { &self.0 } - pub fn canonicalize(&self) -> Option { - CanonicalItemUrl::parse(&self.0) + pub fn canonicalize(&self) -> Option { + ItemId::parse(&self.0) } } @@ -496,7 +152,7 @@ impl RelativePath { &self.0 } - pub fn join_under_ontology_root(&self, root: &CanonicalItemUrl) -> Option { + pub fn join_under_ontology_root(&self, root: &ItemId) -> Option { let base = root.tilde_tail()?; let joined = if base.is_empty() { if self.0.is_empty() { @@ -509,7 +165,7 @@ impl RelativePath { } else { format!("~/{}/{}", base.trim_end_matches('/'), self.0) }; - CanonicalItemUrl::parse(&joined) + ItemId::parse(&joined) } } @@ -545,14 +201,17 @@ impl GardenItemUrl { self.0 } - /// Stored canonical id + RPC `room` field (`"public"` or `"short/slug"`). - pub fn from_stored(stored: &CanonicalItemUrl, room_wire: &str) -> Self { - Self(garden_href_string(stored.as_str(), room_wire)) + /// Stored item id + RPC `room` field (`"public"` or `"short/slug"`). + pub fn from_stored(stored: &ItemId, room_wire: &str) -> Self { + Self(garden_href_string(stored, room_wire)) } /// Like [`Self::from_stored`] but accepts a string that may already be canonical. pub fn from_storage_str(stored: &str, room_wire: &str) -> Self { - Self(garden_href_string(stored, room_wire)) + let Some(id) = ItemId::parse(stored) else { + return Self(api_path_or_url(stored)); + }; + Self(garden_href_string(&id, room_wire)) } } @@ -570,18 +229,15 @@ impl Deref for GardenItemUrl { } } -fn garden_href_string(item: &str, room_wire: &str) -> String { +fn garden_href_string(c: &ItemId, room_wire: &str) -> String { let room = room_wire.trim(); if room.is_empty() || room == "public" { - return api_path_or_url(item); + return api_path_or_url(c.as_str()); } let Some(room_seg) = room_route_segment(room) else { - return api_path_or_url(item); - }; - let Some(c) = CanonicalItemUrl::parse(item) else { - return api_path_or_url(item); + return api_path_or_url(c.as_str()); }; - let root = CanonicalItemUrl::ontology_root(); + let root = ItemId::ontology_root(); let item_norm = c.as_str().trim_end_matches('/'); let root_norm = root.as_str().trim_end_matches('/'); if let Some(tail) = c.tilde_tail() { @@ -600,7 +256,7 @@ fn garden_href_string(item: &str, room_wire: &str) -> String { let tail = tail.strip_prefix("-/").unwrap_or(tail.as_str()); return format!("https://slug.social/r/{room_seg}/-/{tail}"); } - api_path_or_url(item) + api_path_or_url(c.as_str()) } /// Forum thread URL for JSON (`/t/…` or `/r/…/t/…` on slug.social). @@ -650,7 +306,7 @@ impl Deref for ForumThreadUrl { pub struct TildeOntologyPath(pub String); impl TildeOntologyPath { - pub fn from_stored(c: &CanonicalItemUrl) -> Self { + pub fn from_stored(c: &ItemId) -> Self { Self(c.display_path()) } @@ -679,22 +335,22 @@ mod tests { #[test] fn canonical_parent_deep() { - let c = CanonicalItemUrl::parse("~/a/b/c").unwrap(); + let c = ItemId::parse("~/a/b/c").unwrap(); assert_eq!(c.parent().unwrap().as_str(), "https://slug.social/~/a/b"); } #[test] fn canonical_parent_one_level() { - let c = CanonicalItemUrl::parse("~/a").unwrap(); + let c = ItemId::parse("~/a").unwrap(); assert_eq!(c.parent().unwrap().as_str(), "https://slug.social/~"); } #[test] fn canonical_parent_root_is_none() { - let root = CanonicalItemUrl::parse("~/").unwrap(); + let root = ItemId::parse("~/").unwrap(); assert!(root.parent().is_none()); assert_eq!(root.as_str(), SLUG_TILDE_ONTOLOGY_ROOT); - assert_eq!(root, CanonicalItemUrl::ontology_root()); + assert_eq!(root, ItemId::ontology_root()); } #[test] @@ -705,49 +361,49 @@ mod tests { SLUG_TILDE_ONTOLOGY_ROOT.to_string() ); assert_eq!( - CanonicalItemUrl::parse("https://slug.social/~/") + ItemId::parse("https://slug.social/~/") .unwrap() .as_str(), SLUG_TILDE_ONTOLOGY_ROOT ); - let legacy = CanonicalItemUrl("https://slug.social/~/".to_string()); + let legacy = ItemId::opaque("https://slug.social/~/".to_string()); assert_eq!(legacy.normalized_storage().as_str(), SLUG_TILDE_ONTOLOGY_ROOT); } #[test] fn tilde_http_path_tail_maps_router_segment() { assert_eq!( - TildeHttpPathTail::new("").to_canonical(), - CanonicalItemUrl::ontology_root() + TildeHttpPathTail::new("").to_item_id(), + ItemId::ontology_root() ); assert_eq!( - tilde_http_path_to_canonical("topic/x").as_str(), + tilde_http_path_to_item_id("topic/x").as_str(), "https://slug.social/~/topic/x" ); } #[test] fn display_path_slug_ontology_root() { - let r = CanonicalItemUrl::ontology_root(); + let r = ItemId::ontology_root(); assert_eq!(r.display_path(), "~/"); assert_eq!(r.tilde_tail(), Some("")); } #[test] fn tilde_segments_deep() { - let c = CanonicalItemUrl::parse("~/a/b").unwrap(); + let c = ItemId::parse("~/a/b").unwrap(); assert_eq!(c.tilde_segments(), vec!["~", "a", "b"]); } #[test] fn tilde_segments_root() { - let c = CanonicalItemUrl::parse("~/").unwrap(); + let c = ItemId::parse("~/").unwrap(); assert_eq!(c.tilde_segments(), vec!["~"]); } #[test] fn tilde_segments_non_ontology_is_empty() { - let c = CanonicalItemUrl::parse("https://example.com/foo").unwrap(); + let c = ItemId::parse("https://example.com/foo").unwrap(); assert_eq!(c.tilde_segments(), Vec::<&str>::new()); } @@ -831,28 +487,28 @@ mod tests { } #[test] - fn canonical_item_url_parent_external_strips_last_segment() { - let c = CanonicalItemUrl::parse("https://spotify.com/track/1").unwrap(); + fn item_id_parent_external_strips_last_segment() { + let c = ItemId::parse("https://spotify.com/track/1").unwrap(); assert_eq!( c.parent().unwrap().as_str(), "https://spotify.com/track" ); assert_eq!( - CanonicalItemUrl::parse("https://github.com/iss/1") + ItemId::parse("https://github.com/iss/1") .unwrap() .parent() .unwrap() .as_str(), "https://github.com/iss" ); - assert!(CanonicalItemUrl::parse("https://github.com").unwrap().parent().is_none()); + assert!(ItemId::parse("https://github.com").unwrap().parent().is_none()); } #[test] fn display_path_roundtrips_dash_and_tilde() { - let ext = CanonicalItemUrl::parse("https://GitHub.com/org/Issue").unwrap(); + let ext = ItemId::parse("https://GitHub.com/org/Issue").unwrap(); assert_eq!(ext.display_path(), "-/github.com/org/issue"); - let tilde = CanonicalItemUrl::parse("~/Rust/Doc").unwrap(); + let tilde = ItemId::parse("~/Rust/Doc").unwrap(); assert_eq!(tilde.display_path(), "~/rust/doc"); } Side B — contributor: tommy-mor Side B — commit message: [52f5c51c] Add Reddit OAuth linking and make UUID the only account identity. OAuth providers only attach to a session UUID (first link creates the principal); linked providers stay private on the account page. Co-authored-by: Cursor Side B — unified diff (full patch): diff --git a/AGENTS.md b/AGENTS.md index 6e0fd8ebb65d665c9c1438e3275971d62b98fd95..e9cc3173dbeb21ad0fc090ca7b407b027c7820a9 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -35,8 +35,11 @@ Environment variables (defaults in `server/src/state.rs`): - `SORTER2_DATA_DIR` — default `./data` (created on startup) - `SORTER2_EVENT_LOG` — default `{data_dir}/events.jsonl` - `SORTER2_BASE_URL` — public origin (also drives Secure cookies when `https://`) -- `GITHUB_CLIENT_ID` / `GITHUB_CLIENT_SECRET` — GitHub OAuth (optional; login disabled if unset) -- `SORTER2_ALLOW_MOCK_OAUTH=1` — allow `mock_user` on `/auth/github` (tests only) +- `GITHUB_CLIENT_ID` / `GITHUB_CLIENT_SECRET` — GitHub OAuth linking (optional) +- `REDDIT_CLIENT_ID` / `REDDIT_CLIENT_SECRET` (or `REDDIT_APP_*`) — Reddit API import + OAuth linking (optional) +- `SORTER2_ALLOW_MOCK_OAUTH=1` — allow `mock_user` on `/auth/github` and `/auth/reddit` (tests only) + +Identity: UUID is canonical. OAuth providers only *link* to a UUID (first link creates the principal). Linked providers are private to the account owner. Health check: `GET /healthz` → `ok`. diff --git a/server/src/auth/mod.rs b/server/src/auth/mod.rs index 5906f93b13853421e96a3c37bc9d8202a47842bf..c706ae8045a811e5941f5f6c72da88f42a403a82 100644 --- a/server/src/auth/mod.rs +++ b/server/src/auth/mod.rs @@ -1,4 +1,8 @@ -//! GitHub OAuth login, session cookies, and vote actor resolution. +//! OAuth linking, session cookies, and vote actor resolution. +//! +//! Canonical identity is a UUID. OAuth providers only *link* to that UUID +//! (first link creates the principal; later links attach while logged in). +//! Which providers are linked is private to the account owner. pub mod config; pub mod identity; @@ -22,7 +26,9 @@ use crate::{ form_template::template_json_compact, html::layout, state::AppState, - storage_schema::{oauth_link_owner, pseudonym_owner, Store, StoreFields}, + storage_schema::{ + linked_providers_for_uuid, oauth_link_owner, pseudonym_owner, Store, StoreFields, + }, ui_action::UI_RPC_FIELD, }; @@ -53,10 +59,12 @@ fn new_actor_uuid() -> String { pub struct LoginQuery { #[serde(default)] pub return_to: Option, + #[serde(default)] + pub error: Option, } #[derive(Debug, Deserialize)] -pub struct GitHubStartQuery { +pub struct OAuthStartQuery { #[serde(default)] pub return_to: Option, #[serde(default)] @@ -72,15 +80,22 @@ fn return_from_query_or_jar(jar: &CookieJar, query: Option<&str>) -> String { .unwrap_or_else(|| "/".to_string()) } -fn oauth_providers(base_url: &str, return_to: &str) -> Vec<(&'static str, String)> { +/// Available OAuth link targets: `(provider_key, label, start_href)`. +fn oauth_providers(base_url: &str, return_to: &str) -> Vec<(&'static str, &'static str, String)> { let mut out = Vec::new(); + let enc = urlencoding::encode(return_to); if oauth::GitHubConfig::from_env(base_url).is_some() { out.push(( - "GitHub", - format!( - "/auth/github?return_to={}", - urlencoding::encode(return_to) - ), + "github", + oauth::provider_label("github"), + format!("/auth/github?return_to={enc}"), + )); + } + if oauth::RedditConfig::from_env(base_url).is_some() { + out.push(( + "reddit", + oauth::provider_label("reddit"), + format!("/auth/reddit?return_to={enc}"), )); } out @@ -125,23 +140,41 @@ fn alias_claim_forms(return_to: &str, submit_label: &str) -> Result Markup { +fn login_error_message(code: Option<&str>) -> Option<&'static str> { + match code { + Some("oauth_taken") => { + Some("that OAuth account is already linked to a different sorter2 account") + } + Some("oauth_failed") => Some("OAuth failed — try again"), + _ => None, + } +} + +fn signed_out_body( + providers: &[(&str, &str, String)], + error: Option<&str>, +) -> Markup { html! { main class="panel login-page" { section class="login-section" { h1 { "sign in" } - p class="muted" { "link an account to vote under a lasting alias" } + p class="muted" { + "link an OAuth account to create your identity, then claim an alias to vote" + } + @if let Some(msg) = login_error_message(error) { + p class="alias-bad" data-testid="login-error" { (msg) } + } @if providers.is_empty() { p class="muted" { - "OAuth is not configured. Set GITHUB_CLIENT_ID and GITHUB_CLIENT_SECRET." + "OAuth is not configured. Set GitHub and/or Reddit client credentials." } } @else { ul class="oauth-provider-list" { - @for (name, href) in providers { + @for (key, label, href) in providers { li { a href=(href) class="btn-primary oauth-provider" - data-testid=(format!("oauth-{}", name.to_lowercase())) { - (format!("Continue with {name}")) + data-testid=(format!("oauth-{key}")) { + (format!("Link {label}")) } } } @@ -156,7 +189,10 @@ fn signed_out_body(providers: &[(&str, String)]) -> Markup { fn account_body( actor: &session::SessionActor, aliases: &[String], - providers: &[(&str, String)], + // Provider keys already linked to this UUID (private). + linked: &[String], + // Providers available to link: not yet attached. + unlinkable: &[(&str, &str, String)], claim_forms: Markup, ) -> Markup { let current = actor.pseudonym.trim(); @@ -212,16 +248,29 @@ fn account_body( (claim_forms) } - @if !providers.is_empty() { - section class="login-section" { - h2 { "linked sign-in" } - p class="muted small" { "sign in again with the same provider to return to this account" } + section class="login-section" { + h2 { "linked sign-in" } + p class="muted small" { + "private to you — linking more providers raises trust weight without publishing which accounts you use" + } + @if linked.is_empty() { + p class="muted" data-testid="linked-providers-empty" { "none yet" } + } @else { + ul class="linked-provider-list" data-testid="linked-providers" { + @for key in linked { + li data-testid=(format!("linked-{key}")) { + (oauth::provider_label(key)) + } + } + } + } + @if !unlinkable.is_empty() { ul class="oauth-provider-list" { - @for (name, href) in providers { + @for (key, label, href) in unlinkable { li { a href=(href) class="btn-secondary oauth-provider" - data-testid=(format!("oauth-relink-{}", name.to_lowercase())) { - (format!("Re-link {name}")) + data-testid=(format!("oauth-link-{key}")) { + (format!("Link {label}")) } } } @@ -243,12 +292,21 @@ fn account_body( fn login_body( session: Option<&session::SessionActor>, aliases: &[String], - providers: &[(&str, String)], + linked: &[String], + providers: &[(&str, &str, String)], claim_forms: Option, + error: Option<&str>, ) -> Markup { match (session, claim_forms) { - (Some(actor), Some(forms)) => account_body(actor, aliases, providers, forms), - _ => signed_out_body(providers), + (Some(actor), Some(forms)) => { + let unlinkable: Vec<_> = providers + .iter() + .filter(|(key, _, _)| !linked.iter().any(|p| p == key)) + .cloned() + .collect(); + account_body(actor, aliases, linked, &unlinkable, forms) + } + _ => signed_out_body(providers, error), } } @@ -268,6 +326,10 @@ pub async fn login_page( .as_ref() .map(|s| alias_list(db, &s.uuid)) .unwrap_or_default(); + let linked = session + .as_ref() + .map(|s| linked_providers_for_uuid(db, &s.uuid).unwrap_or_default()) + .unwrap_or_default(); let providers = oauth_providers(&base_url_from_env(state.cfg.port), &return_to); let claim_forms = if session.is_some() { @@ -282,7 +344,14 @@ pub async fn login_page( } else { "login · sorter2" }, - login_body(session.as_ref(), &aliases, &providers, claim_forms), + login_body( + session.as_ref(), + &aliases, + &linked, + &providers, + claim_forms, + query.error.as_deref(), + ), state.views.get_views("/login"), session .as_ref() @@ -302,7 +371,6 @@ pub async fn alias_page( let db = state.projection_store.db(); let session = session::load_valid_session(db, &session_id).ok_or(StatusCode::UNAUTHORIZED)?; if session::session_has_pseudonym(&session) { - // Already onboarded — manage aliases on the account page. return Ok(Redirect::to("/login").into_response()); } @@ -331,7 +399,7 @@ pub async fn alias_page( pub async fn github_start( State(state): State, jar: CookieJar, - Query(query): Query, + Query(query): Query, ) -> Result { let cfg = oauth::GitHubConfig::from_env(&base_url_from_env(state.cfg.port)) .ok_or(StatusCode::SERVICE_UNAVAILABLE)?; @@ -342,7 +410,28 @@ pub async fn github_start( } else { None }; - let url = oauth::authorize_url(&cfg, &state_token, mock_user); + let url = oauth::github_authorize_url(&cfg, &state_token, mock_user); + let jar = jar + .add(session::oauth_state_cookie_value(&state_token)) + .add(session::auth_return_cookie_value(&return_to)); + Ok((jar, Redirect::temporary(&url)).into_response()) +} + +pub async fn reddit_start( + State(state): State, + jar: CookieJar, + Query(query): Query, +) -> Result { + let cfg = oauth::RedditConfig::from_env(&base_url_from_env(state.cfg.port)) + .ok_or(StatusCode::SERVICE_UNAVAILABLE)?; + let return_to = return_from_query_or_jar(&jar, query.return_to.as_deref()); + let state_token = session::new_oauth_state(); + let mock_user = if config::mock_oauth_allowed() { + query.mock_user.as_deref() + } else { + None + }; + let url = oauth::reddit_authorize_url(&cfg, &state_token, mock_user); let jar = jar .add(session::oauth_state_cookie_value(&state_token)) .add(session::auth_return_cookie_value(&return_to)); @@ -355,6 +444,13 @@ pub struct OAuthCallbackQuery { pub state: String, } +/// Link `provider:provider_id` to a UUID. +/// +/// - Logged in + new provider → attach to session UUID +/// - Logged in + already ours → no-op +/// - Logged in + owned by someone else → conflict +/// - Logged out + known link → resume that UUID +/// - Logged out + unknown → create principal + first link async fn finish_oauth_login( state: &AppState, jar: CookieJar, @@ -364,11 +460,41 @@ async fn finish_oauth_login( let db = state.projection_store.db(); let return_to = return_from_query_or_jar(&jar, None); - let uuid = match oauth_link_owner(db, provider, &provider_id) - .map_err(|_| StatusCode::INTERNAL_SERVER_ERROR)? - { - Some(existing) => existing, - None => { + let existing_owner = oauth_link_owner(db, provider, &provider_id) + .map_err(|_| StatusCode::INTERNAL_SERVER_ERROR)?; + + let session_uuid = session::session_id_from_jar(&jar) + .as_deref() + .and_then(|id| session::load_valid_session(db, id)) + .map(|s| s.uuid); + let linking_while_logged_in = session_uuid.is_some(); + + let uuid = match (session_uuid, existing_owner) { + (Some(session_uuid), Some(owner)) if owner == session_uuid => session_uuid, + (Some(_), Some(_)) => { + return Ok(( + jar.add(session::clear_oauth_state_cookie()), + "/login?error=oauth_taken".into(), + )); + } + (Some(session_uuid), None) => { + let ts = now_ms(); + state + .append_identity_events(vec![Event::OauthLinked { + uuid: session_uuid.clone(), + provider: provider.to_string(), + provider_id, + ts, + }]) + .await + .map_err(|e| { + tracing::warn!(err = %e, "oauth link append failed"); + StatusCode::INTERNAL_SERVER_ERROR + })?; + session_uuid + } + (None, Some(owner)) => owner, + (None, None) => { let uuid = new_actor_uuid(); let ts = now_ms(); state @@ -409,6 +535,9 @@ async fn finish_oauth_login( "/login/alias?return_to={}", urlencoding::encode(&return_to) ) + } else if linking_while_logged_in { + // Additional link while already in an account → stay on account page. + "/login".to_string() } else { return_to }; @@ -434,22 +563,57 @@ pub async fn github_callback( .build() .map_err(|_| StatusCode::INTERNAL_SERVER_ERROR)?; - let token = oauth::exchange_code(&client, &cfg, &query.code) + let token = oauth::github_exchange_code(&client, &cfg, &query.code) .await .map_err(|e| { tracing::warn!(err = %e, "github oauth token exchange failed"); StatusCode::BAD_GATEWAY })?; - let user = oauth::fetch_user(&client, &cfg.api_base, &token) + let user = oauth::github_fetch_user(&client, &cfg.api_base, &token) .await .map_err(|e| { tracing::warn!(err = %e, "github user fetch failed"); StatusCode::BAD_GATEWAY })?; - let provider = "github"; - let provider_id = oauth::provider_id(&user); - let (jar, dest) = finish_oauth_login(&state, jar, provider, provider_id).await?; + let (jar, dest) = + finish_oauth_login(&state, jar, "github", oauth::github_provider_id(&user)).await?; + Ok((jar, Redirect::to(&dest)).into_response()) +} + +pub async fn reddit_callback( + State(state): State, + jar: CookieJar, + Query(query): Query, +) -> Result { + let cfg = oauth::RedditConfig::from_env(&base_url_from_env(state.cfg.port)) + .ok_or(StatusCode::SERVICE_UNAVAILABLE)?; + + let expected_state = session::oauth_state_from_jar(&jar).ok_or(StatusCode::BAD_REQUEST)?; + if expected_state != query.state { + return Err(StatusCode::BAD_REQUEST); + } + + let client = Client::builder() + .timeout(std::time::Duration::from_secs(15)) + .build() + .map_err(|_| StatusCode::INTERNAL_SERVER_ERROR)?; + + let token = oauth::reddit_exchange_code(&client, &cfg, &query.code) + .await + .map_err(|e| { + tracing::warn!(err = %e, "reddit oauth token exchange failed"); + StatusCode::BAD_GATEWAY + })?; + let user = oauth::reddit_fetch_user(&client, &cfg, &token) + .await + .map_err(|e| { + tracing::warn!(err = %e, "reddit user fetch failed"); + StatusCode::BAD_GATEWAY + })?; + + let (jar, dest) = + finish_oauth_login(&state, jar, "reddit", oauth::reddit_provider_id(&user)).await?; Ok((jar, Redirect::to(&dest)).into_response()) } diff --git a/server/src/auth/oauth.rs b/server/src/auth/oauth.rs index b80078fee0c5acd905b9125d29454e46c4e0d066..369b1527d9032cf312a07819279e0f988ef4a36e 100644 --- a/server/src/auth/oauth.rs +++ b/server/src/auth/oauth.rs @@ -1,8 +1,15 @@ -//! GitHub OAuth (raw reqwest, same style as reddit.rs). +//! OAuth providers (GitHub + Reddit). Provider accounts only *link* to a UUID; +//! the UUID is the canonical identity. Which providers are linked is private. use reqwest::Client; use serde::Deserialize; +use crate::reddit::{ + default_user_agent, reddit_oauth_api_base, reddit_oauth_token_base, +}; + +// ── GitHub ────────────────────────────────────────────────────────────────── + #[derive(Debug, Clone)] pub struct GitHubConfig { pub client_id: String, @@ -40,7 +47,7 @@ impl GitHubConfig { } #[derive(Debug, Deserialize)] -struct TokenResponse { +struct GitHubTokenResponse { access_token: String, } @@ -50,7 +57,7 @@ pub struct GitHubUser { pub login: String, } -pub fn authorize_url(cfg: &GitHubConfig, state: &str, mock_user: Option<&str>) -> String { +pub fn github_authorize_url(cfg: &GitHubConfig, state: &str, mock_user: Option<&str>) -> String { let mut url = format!( "{}/login/oauth/authorize?client_id={}&redirect_uri={}&scope=read:user&state={}", cfg.oauth_base.trim_end_matches('/'), @@ -65,7 +72,7 @@ pub fn authorize_url(cfg: &GitHubConfig, state: &str, mock_user: Option<&str>) - url } -pub async fn exchange_code( +pub async fn github_exchange_code( client: &Client, cfg: &GitHubConfig, code: &str, @@ -90,14 +97,14 @@ pub async fn exchange_code( return Err(format!("github token HTTP {}", resp.status())); } - let body: TokenResponse = resp + let body: GitHubTokenResponse = resp .json() .await .map_err(|e| format!("github token parse failed: {e}"))?; Ok(body.access_token) } -pub async fn fetch_user( +pub async fn github_fetch_user( client: &Client, api_base: &str, access_token: &str, @@ -120,10 +127,149 @@ pub async fn fetch_user( .map_err(|e| format!("github user parse failed: {e}")) } -pub fn provider_id(user: &GitHubUser) -> String { +pub fn github_provider_id(user: &GitHubUser) -> String { user.id.to_string() } +// ── Reddit ────────────────────────────────────────────────────────────────── + +#[derive(Debug, Clone)] +pub struct RedditConfig { + pub client_id: String, + pub client_secret: String, + pub redirect_uri: String, + /// Host for `/api/v1/authorize` (www.reddit.com in production). + pub authorize_base: String, + /// Host for `POST /api/v1/access_token`. + pub token_base: String, + /// Host for bearer `GET /api/v1/me` (oauth.reddit.com). + pub api_base: String, + pub user_agent: String, +} + +/// Authorize page base; defaults to the same host as token POSTs. +pub fn reddit_authorize_base() -> String { + std::env::var("REDDIT_OAUTH_AUTHORIZE_BASE") + .or_else(|_| std::env::var("REDDIT_OAUTH_BASE")) + .unwrap_or_else(|_| "https://www.reddit.com".into()) +} + +impl RedditConfig { + pub fn from_env(base_url: &str) -> Option { + let client_id = std::env::var("REDDIT_CLIENT_ID") + .or_else(|_| std::env::var("REDDIT_APP_ID")) + .ok()?; + let client_secret = std::env::var("REDDIT_CLIENT_SECRET") + .or_else(|_| std::env::var("REDDIT_APP_SECRET")) + .ok()?; + if client_id.is_empty() || client_secret.is_empty() { + return None; + } + let base = base_url.trim_end_matches('/'); + Some(Self { + client_id, + client_secret, + redirect_uri: format!("{base}/auth/reddit/callback"), + authorize_base: reddit_authorize_base(), + token_base: reddit_oauth_token_base(), + api_base: reddit_oauth_api_base(), + user_agent: default_user_agent(), + }) + } +} + +#[derive(Debug, Deserialize)] +struct RedditTokenResponse { + access_token: String, +} + +#[derive(Debug, Deserialize)] +pub struct RedditUser { + /// Stable id (`t2_…`); never use `name` as identity. + pub id: String, + pub name: String, +} + +pub fn reddit_authorize_url(cfg: &RedditConfig, state: &str, mock_user: Option<&str>) -> String { + let mut url = format!( + "{}/api/v1/authorize?client_id={}&response_type=code&state={}&redirect_uri={}&duration=temporary&scope=identity", + cfg.authorize_base.trim_end_matches('/'), + urlencoding::encode(&cfg.client_id), + urlencoding::encode(state), + urlencoding::encode(&cfg.redirect_uri), + ); + if let Some(user) = mock_user { + url.push_str("&mock_user="); + url.push_str(&urlencoding::encode(user)); + } + url +} + +pub async fn reddit_exchange_code( + client: &Client, + cfg: &RedditConfig, + code: &str, +) -> Result { + let resp = client + .post(format!( + "{}/api/v1/access_token", + cfg.token_base.trim_end_matches('/') + )) + .header("User-Agent", &cfg.user_agent) + .basic_auth(&cfg.client_id, Some(&cfg.client_secret)) + .form(&[ + ("grant_type", "authorization_code"), + ("code", code), + ("redirect_uri", cfg.redirect_uri.as_str()), + ]) + .send() + .await + .map_err(|e| format!("reddit token request failed: {e}"))?; + + if !resp.status().is_success() { + let status = resp.status(); + let body = resp.text().await.unwrap_or_default(); + return Err(format!("reddit token HTTP {status}: {body}")); + } + + let body: RedditTokenResponse = resp + .json() + .await + .map_err(|e| format!("reddit token parse failed: {e}"))?; + Ok(body.access_token) +} + +pub async fn reddit_fetch_user( + client: &Client, + cfg: &RedditConfig, + access_token: &str, +) -> Result { + let resp = client + .get(format!( + "{}/api/v1/me", + cfg.api_base.trim_end_matches('/') + )) + .header("User-Agent", &cfg.user_agent) + .bearer_auth(access_token) + .send() + .await + .map_err(|e| format!("reddit user request failed: {e}"))?; + + if !resp.status().is_success() { + return Err(format!("reddit user HTTP {}", resp.status())); + } + + resp.json() + .await + .map_err(|e| format!("reddit user parse failed: {e}")) +} + +pub fn reddit_provider_id(user: &RedditUser) -> String { + user.id.clone() +} + +// ── Shared helpers ────────────────────────────────────────────────────────── + pub fn validate_pseudonym(raw: &str) -> Result { let trimmed = raw.trim(); if trimmed.is_empty() { @@ -144,3 +290,12 @@ pub fn validate_pseudonym(raw: &str) -> Result { pub fn sanitize_pseudonym(login: &str) -> String { validate_pseudonym(login).unwrap_or_else(|_| "user".to_string()) } + +/// Display name for a provider key (`github` → `GitHub`). Never show provider ids. +pub fn provider_label(provider: &str) -> &'static str { + match provider { + "github" => "GitHub", + "reddit" => "Reddit", + _ => "OAuth", + } +} diff --git a/server/src/lib.rs b/server/src/lib.rs index 84f2565b105fb302241b949af64bd5e49916eab2..0d948af1457504fdbd34d0b14261f65ac58da0ec 100644 --- a/server/src/lib.rs +++ b/server/src/lib.rs @@ -47,6 +47,8 @@ pub fn create_app(state: AppState) -> Router { .route("/login/alias", get(crate::auth::alias_page)) .route("/auth/github", get(crate::auth::github_start)) .route("/auth/github/callback", get(crate::auth::github_callback)) + .route("/auth/reddit", get(crate::auth::reddit_start)) + .route("/auth/reddit/callback", get(crate::auth::reddit_callback)) .route("/auth/logout", post(crate::auth::logout)) .route("/auth/switch", post(crate::auth::switch_pseudonym)) .route("/ui", post(crate::api::ui_html::post_ui_html)) diff --git a/server/src/projection_apply.rs b/server/src/projection_apply.rs index 4fb4b48ee0a53b7f673fae0c9731eed5acc33332..f50458c2ff3c447a3cfd28adb298e6883a2da4e9 100644 --- a/server/src/projection_apply.rs +++ b/server/src/projection_apply.rs @@ -4,6 +4,8 @@ //! child links, recent-vote appends) plus a cursor advance, all committed in //! one atomic `DisableWal` batch. +use std::collections::HashMap; + use crate::{ event_log::EventLogError, events::{Event, EventRecord}, @@ -44,6 +46,8 @@ pub fn apply_records( let db = projection_store.db(); let mut batch = db.batch(); let mut last_seq = 0u64; + // Weight reads must see earlier writes in this same batch. + let mut pending_weights: HashMap = HashMap::new(); for record in records { match &record.event { @@ -85,6 +89,7 @@ pub fn apply_records( ensure_path_writes(&mut batch, &parsed); } Event::PrincipalCreated { uuid, .. } => { + pending_weights.insert(uuid.clone(), BASE_TRUST_WEIGHT); batch.write( Store::root() .user_weights() @@ -112,17 +117,25 @@ pub fn apply_records( } } else { batch.write(Store::root().oauth_links().key(&link_key).set(uuid)); - let current = Store::root() - .user_weights() - .key(&uuid.clone()) - .get(db) - .map_err(|e| EventLogError::Apply(e.to_string()))? + let current = pending_weights + .get(uuid) + .copied() + .or_else(|| { + Store::root() + .user_weights() + .key(&uuid.clone()) + .get(db) + .ok() + .flatten() + }) .unwrap_or(BASE_TRUST_WEIGHT); + let next = trust_weight_after_link(current); + pending_weights.insert(uuid.clone(), next); batch.write( Store::root() .user_weights() .key(&uuid.clone()) - .set(&trust_weight_after_link(current)), + .set(&next), ); } } @@ -205,6 +218,15 @@ mod tests { ), record( 3, + Event::OauthLinked { + uuid: uuid.into(), + provider: "reddit".into(), + provider_id: "t2_abc".into(), + ts, + }, + ), + record( + 4, Event::PseudonymClaimed { uuid: uuid.into(), pseudonym: "octocat".into(), @@ -219,8 +241,12 @@ mod tests { oauth_link_owner(store.db(), "github", "42").unwrap(), Some(uuid.to_string()) ); + assert_eq!( + crate::storage_schema::linked_providers_for_uuid(store.db(), uuid).unwrap(), + vec!["github".to_string(), "reddit".to_string()] + ); assert_eq!(resolve_actor_uuid(store.db(), "octocat").unwrap(), uuid); - assert_eq!(user_trust_weight(store.db(), uuid).unwrap(), 1.5); + assert_eq!(user_trust_weight(store.db(), uuid).unwrap(), 2.0); let aliases = Store::root() .user_pseudonyms() .key(&uuid.to_string()) diff --git a/server/src/storage_schema.rs b/server/src/storage_schema.rs index b942810afd96d3bf01b7765718303a39be4ef8d2..dac6f8e6080e3f5683c24dbb8861f5a981d456c4 100644 --- a/server/src/storage_schema.rs +++ b/server/src/storage_schema.rs @@ -103,6 +103,25 @@ pub fn oauth_link_owner(db: &Db, provider: &str, provider_id: &str) -> durable:: .get(db) } +/// Provider names linked to a UUID (`github`, `reddit`, …). Private — for the +/// account owner's page only; never expose which providers are linked publicly. +pub fn linked_providers_for_uuid(db: &Db, uuid: &str) -> durable::Result> { + let mut providers = Vec::new(); + for (key, owner) in Store::root().oauth_links().iter(db)? { + if owner != uuid { + continue; + } + let Some((provider, _)) = key.split_once(':') else { + continue; + }; + if !providers.iter().any(|p| p == provider) { + providers.push(provider.to_string()); + } + } + providers.sort(); + Ok(providers) +} + pub const RECENT_VOTES_CAP: u64 = 200; fn id_key(id: &ItemId) -> String { diff --git a/test/support/harness.clj b/test/support/harness.clj index 4505ece5aa193e826ca61c52f1467d252080d66b..741024aab1d14269e225791189633287980ae2e4 100644 --- a/test/support/harness.clj +++ b/test/support/harness.clj @@ -48,12 +48,15 @@ "GITHUB_CLIENT_SECRET" "test-secret" "GITHUB_OAUTH_BASE" (str "http://127.0.0.1:" oauth-port) "GITHUB_API_BASE" (str "http://127.0.0.1:" oauth-port) + ;; Reddit import fixtures + Reddit OAuth on reddit-port. "REDDIT_API_BASE" (str "http://127.0.0.1:" reddit-port) + "REDDIT_CLIENT_ID" "test-reddit" + "REDDIT_CLIENT_SECRET" "test-reddit-secret" "REDDIT_OAUTH_BASE" (str "http://127.0.0.1:" reddit-port) - "REDDIT_CLIENT_ID" "" - "REDDIT_CLIENT_SECRET" "" - "REDDIT_APP_ID" "" - "REDDIT_APP_SECRET" ""})) + "REDDIT_OAUTH_AUTHORIZE_BASE" (str "http://127.0.0.1:" reddit-port) + "REDDIT_OAUTH_TOKEN_BASE" (str "http://127.0.0.1:" reddit-port) + "REDDIT_OAUTH_API_BASE" (str "http://127.0.0.1:" reddit-port) + "REDDIT_USER_AGENT" "web:sorter2-test:v0 (by /u/test)"})) (defn with-auth-servers "Start mock Reddit + mock OAuth + release sorter2-server. diff --git a/test/support/mock_oauth.clj b/test/support/mock_oauth.clj index 909d7a7be8b159ea54a122d69c46af3db3318118..5ba7e9be3cf6d2226648e9609a09ed45f306930f 100644 --- a/test/support/mock_oauth.clj +++ b/test/support/mock_oauth.clj @@ -1,5 +1,5 @@ (ns test.support.mock-oauth - "In-process HTTP stub for GitHub OAuth (authorize, token, /user)." + "In-process HTTP stub for GitHub + Reddit OAuth (authorize, token, user)." (:require [clojure.string :as str]) (:import [com.sun.net.httpserver HttpServer HttpHandler HttpExchange] [java.net InetSocketAddress URLDecoder])) @@ -12,11 +12,14 @@ (URLDecoder/decode (or v "") "UTF-8")))) (str/split query #"&")))) -(defn- parse-mock-user [raw] +(defn- parse-mock-user + "GitHub-style `id:login` (numeric id). Reddit-style `t2_xxx:name`." + [raw] (let [s (or raw "1002:newbie") [id login] (str/split s #":" 2)] - {:id (Long/parseLong id) - :login (or login "newbie")})) + {:id id + :login (or login "newbie") + :numeric? (re-matches #"\d+" id)})) (defn- send-json [^HttpExchange ex status body] (let [bytes (.getBytes body "UTF-8")] @@ -31,9 +34,10 @@ (.sendResponseHeaders ex 302 -1) (.close (.getResponseBody ex))) -(defn- read-form-code [^HttpExchange ex] +(defn- read-form [^HttpExchange ex] (let [body (slurp (.getInputStream ex))] - (query-param body "code"))) + {:code (query-param body "code") + :grant (query-param body "grant_type")})) (defn- bearer-token [^HttpExchange ex] (some-> (.getRequestHeaders ex) @@ -44,8 +48,18 @@ (when (str/starts-with? token "mock:") (parse-mock-user (subs token 5)))) +(defn- authorize-redirect [exchange query] + (let [redirect-uri (query-param query "redirect_uri") + state (query-param query "state") + mock-user (query-param query "mock_user") + user (parse-mock-user mock-user) + code (str "mock:" (:id user) ":" (:login user)) + loc (str redirect-uri "?code=" (java.net.URLEncoder/encode code "UTF-8") + "&state=" (java.net.URLEncoder/encode state "UTF-8"))] + (send-redirect exchange loc))) + (defn start-mock-oauth - "Start mock GitHub OAuth on `port`. Returns a zero-arg `stop` function." + "Start mock GitHub + Reddit OAuth on `port`. Returns a zero-arg `stop` function." [port] (let [server (HttpServer/create (InetSocketAddress. "127.0.0.1" port) 0) handler @@ -53,28 +67,45 @@ (handle [^HttpExchange exchange] (let [uri (.getRequestURI exchange) path (.getPath uri) - query (.getQuery uri)] + query (.getQuery uri) + method (.getRequestMethod exchange)] (cond + ;; GitHub authorize (str/ends-with? path "/login/oauth/authorize") - (let [redirect-uri (query-param query "redirect_uri") - state (query-param query "state") - mock-user (query-param query "mock_user") - user (parse-mock-user mock-user) - code (str "mock:" (:id user) ":" (:login user)) - loc (str redirect-uri "?code=" (java.net.URLEncoder/encode code "UTF-8") - "&state=" (java.net.URLEncoder/encode state "UTF-8"))] - (send-redirect exchange loc)) + (authorize-redirect exchange query) + + ;; Reddit authorize + (str/ends-with? path "/api/v1/authorize") + (authorize-redirect exchange query) - (str/ends-with? path "/login/oauth/access_token") - (let [code (or (read-form-code exchange) "mock:1002:newbie")] + ;; GitHub token + (and (= method "POST") (str/ends-with? path "/login/oauth/access_token")) + (let [code (or (:code (read-form exchange)) "mock:1002:newbie")] (send-json exchange 200 (str "{\"access_token\":\"" code "\",\"token_type\":\"bearer\"}"))) + ;; Reddit token (client_credentials for import + authorization_code for login) + (and (= method "POST") (str/ends-with? path "/api/v1/access_token")) + (let [form (read-form exchange) + grant (or (:grant form) "") + code (or (:code form) "mock:t2_test:redditor")] + (if (= grant "client_credentials") + (send-json exchange 200 "{\"access_token\":\"app-token\",\"token_type\":\"bearer\",\"expires_in\":3600}") + (send-json exchange 200 (str "{\"access_token\":\"" code "\",\"token_type\":\"bearer\",\"expires_in\":3600}")))) + + ;; GitHub user (= path "/user") (let [token (bearer-token exchange) - user (or (parse-token-user token) {:id 1002 :login "newbie"})] + user (or (parse-token-user token) {:id "1002" :login "newbie" :numeric? true})] (send-json exchange 200 (str "{\"id\":" (:id user) ",\"login\":\"" (:login user) "\"}"))) + ;; Reddit /api/v1/me + (str/ends-with? path "/api/v1/me") + (let [token (bearer-token exchange) + user (or (parse-token-user token) {:id "t2_test" :login "redditor"})] + (send-json exchange 200 + (str "{\"id\":\"" (:id user) "\",\"name\":\"" (:login user) "\"}"))) + :else (send-json exchange 404 "{\"error\":\"not found\"}")))))] (.createContext server "/" handler) diff --git a/test/support/mock_reddit.clj b/test/support/mock_reddit.clj index 5efa92db3e1f79b8f423a2f1adcbda959123c1ad..a630cf0938722193e9af88382d60e777ff371be4 100644 --- a/test/support/mock_reddit.clj +++ b/test/support/mock_reddit.clj @@ -1,14 +1,56 @@ (ns test.support.mock-reddit - "In-process HTTP stub for Reddit API fixtures (`test/fixtures/reddit/`)." + "In-process HTTP stub for Reddit API fixtures + OAuth login endpoints." (:require [clojure.java.io :as io] [clojure.string :as str]) (:import [com.sun.net.httpserver HttpServer HttpHandler HttpExchange] - [java.net InetSocketAddress])) + [java.net InetSocketAddress URLDecoder])) (defn fixtures-dir ([] (fixtures-dir (System/getProperty "user.dir"))) ([root] (str root "/test/fixtures/reddit"))) +(defn- query-param [query key] + (when query + (some (fn [pair] + (let [[k v] (str/split pair "=" 2)] + (when (= k key) + (URLDecoder/decode (or v "") "UTF-8")))) + (str/split query #"&")))) + +(defn- parse-mock-user [raw] + (let [s (or raw "t2_test:redditor") + [id login] (str/split s #":" 2)] + {:id id :login (or login "redditor")})) + +(defn- send-bytes [^HttpExchange ex status ^bytes body content-type] + (.set (.getResponseHeaders ex) "Content-Type" content-type) + (.sendResponseHeaders ex status (alength body)) + (doto (.getResponseBody ex) + (.write body) + (.close))) + +(defn- send-json [^HttpExchange ex status body] + (send-bytes ex status (.getBytes body "UTF-8") "application/json")) + +(defn- send-redirect [^HttpExchange ex location] + (.set (.getResponseHeaders ex) "Location" location) + (.sendResponseHeaders ex 302 -1) + (.close (.getResponseBody ex))) + +(defn- read-form [^HttpExchange ex] + (let [body (slurp (.getInputStream ex))] + {:code (query-param body "code") + :grant (query-param body "grant_type")})) + +(defn- bearer-token [^HttpExchange ex] + (some-> (.getRequestHeaders ex) + (.getFirst "Authorization") + (str/replace #"^[Bb]earer " ""))) + +(defn- parse-token-user [token] + (when (str/starts-with? token "mock:") + (parse-mock-user (subs token 5)))) + (defn start-mock-reddit "Start a mock Reddit API on `port`. Returns a zero-arg `stop` function." ([port] (start-mock-reddit port (fixtures-dir))) @@ -19,15 +61,38 @@ handler (proxy [HttpHandler] [] (handle [^HttpExchange exchange] - ;; `/r//about.json` → subreddit entity; `/r/.json` → listing. - (let [path (.getPath (.getRequestURI exchange)) - body (if (str/includes? path "/about") - about - listing)] - (.sendResponseHeaders exchange 200 (alength body)) - (let [out (.getResponseBody exchange)] - (.write out body) - (.close out)))))] + (let [uri (.getRequestURI exchange) + path (.getPath uri) + query (.getQuery uri) + method (.getRequestMethod exchange)] + (cond + (str/ends-with? path "/api/v1/authorize") + (let [redirect-uri (query-param query "redirect_uri") + state (query-param query "state") + user (parse-mock-user (query-param query "mock_user")) + code (str "mock:" (:id user) ":" (:login user)) + loc (str redirect-uri "?code=" (java.net.URLEncoder/encode code "UTF-8") + "&state=" (java.net.URLEncoder/encode state "UTF-8"))] + (send-redirect exchange loc)) + + (and (= method "POST") (str/ends-with? path "/api/v1/access_token")) + (let [form (read-form exchange) + grant (or (:grant form) "") + code (or (:code form) "mock:t2_test:redditor")] + (if (= grant "client_credentials") + (send-json exchange 200 "{\"access_token\":\"app-token\",\"token_type\":\"bearer\",\"expires_in\":3600}") + (send-json exchange 200 (str "{\"access_token\":\"" code "\",\"token_type\":\"bearer\",\"expires_in\":3600}")))) + + (str/ends-with? path "/api/v1/me") + (let [user (or (parse-token-user (bearer-token exchange)) + {:id "t2_test" :login "redditor"})] + (send-json exchange 200 + (str "{\"id\":\"" (:id user) "\",\"name\":\"" (:login user) "\"}"))) + + ;; `/r//about.json` → subreddit entity; `/r/.json` → listing. + :else + (let [body (if (str/includes? path "/about") about listing)] + (send-bytes exchange 200 body "application/json"))))))] (.createContext server "/" handler) (.setExecutor server nil) (.start server)