You are a constitutional council ranking individual git commits for ownership allocation. Compare these two commits. Decide which contributed more lasting value to the project. Judge substance, not spectacle: - Prefer correct, lasting design and real bugfixes over churn, formatting, renames, or generated noise. - Prefer clarity and necessity over sheer line count. A small precise change can beat a large diffuse one. - Do not favor a side merely because its patch is longer or noisier. - Weight what the change does for the project, not the contributor's name. Return ONLY a JSON object: {"winner": "A" or "B", "ratio": "N:M", "explanation": "..."} The explanation must cite concrete differences in the patches (1-3 sentences). Side A — contributor: tommy-mor Side A — commit message: [1c914c6e] stage set Side A — unified diff (full patch): diff --git a/plan.md b/plan.md new file mode 100644 index 0000000000000000000000000000000000000000..00d6867a1e0ed144a16a020ea037f685ce646c73 --- /dev/null +++ b/plan.md @@ -0,0 +1,155 @@ +# Plan: `ItemId` + `RouteContext` (identity vs hrefs) + +This document is for **the next agent** to continue the refactor without re-deriving context from chat. It supersedes ad-hoc notes: treat it as the checklist of record until the work lands and this file is deleted or trimmed. + +## Goal + +- **Identity** (what lives in the reducer graph, votes, indexes) becomes a **structural `ItemId` enum** in `slug-types`, not a canonical `String` / `CanonicalItemUrl` newtype. +- **Presentation** (tilde / dash display, breadcrumbs) derives from `ItemId` via explicit methods, not string stripping. +- **Routing** (browser `href`s for public vs room) goes through **`RouteContext`** (started in `server/src/html/routing.rs`) so Maud/handlers do not stitch `/r/…` vs `/~` ad hoc. + +**Non-goals for v1 of the migration:** backward-compatible JSONL or dual-read of old canonical strings in the event log (project has accepted breaking changes). If you reintroduce compat, document it here. + +## Current state (as of this plan) + +- **`CanonicalItemUrl`** (`types/src/paths.rs`): newtype around `String`; `parse` / `parent` / `display_path` / `tilde_tail` / etc. Reducer `ContentState`, `VoteData`, ranking, RPC, search, garden, breadcrumbs all use it or `String` keys derived from it. +- **`ThreadNav`** (`server/src/html/forum/nav.rs`): encodes scope prefixes for threads and garden URLs; **`RouteContext`** now wraps `ThreadNav` (`server/src/html/routing.rs`, re-exported from `server/src/html/mod.rs`) but **most HTML still takes `&ThreadNav` directly** — migration incomplete. +- **URL normalization** lives in `types/src/url_normalize.rs` + `canonicalize_item` / `finalize_external_identity_url` in `paths.rs` (YouTube, sorted query params, room path `room_route_segment` in `paths.rs`). +- **Room HTTP paths** are `/r/{short}{slug}` (fused segment); wire **`room_id`** remains `short/slug` for RPC/events. + +## Target architecture + +### `ItemId` (types) + +Suggested shape (adjust after profiling `Ord` / `Hash` / serde size): + +```text +ItemId::Root — tilde ontology root (today `SLUG_TILDE_ONTOLOGY_ROOT`) +ItemId::Local { segments } — slug.social ~/… path as Vec (lowercase segments, non-empty for non-root) +ItemId::External { url: Url } — normalized `url::Url` (crate `url` already in `slug-types`) +``` + +**API surface (minimum):** + +- `ItemId::parse(&str) -> Option` — single entry from DSL / user input / legacy wire (internally may call `canonicalize_item` + structured split). +- `ItemId::to_wire_url(&self) -> String` — only for **external** boundaries if needed (HTTP fetch, rare assertions); avoid using as the primary key once maps use `ItemId`. +- `parent`, `display_path`, `tilde_tail` / `tilde_http_tail`, `tilde_segments`, `last_segment`, `normalized_storage` — port from `CanonicalItemUrl`. +- **`Ord` + `Hash` + `Eq`** stable for `BTreeSet` / `HashMap` (see `write_actor` scope-rank snapshots). +- **`Serialize` / `Deserialize`** — decide **tagged JSON** for any persisted or API-carried structs (e.g. `VoteData` in tests). If RPC must stay stringy for clients, use a **DTO layer** that converts `ItemId` ↔ wire at the boundary only. + +**Remove:** `CanonicalItemUrl` type and all `path_types::CanonicalItemUrl` / `slug_types::paths::CanonicalItemUrl` exports once call sites are migrated. **`Borrow`** on the old newtype goes away; update `nav!` / any code that assumed map keys borrowed as `str`. + +### `RouteContext` (server HTML) + +- **File:** `server/src/html/routing.rs` — **`RouteContext(ThreadNav)`** with `item_href`, `item_href_raw`, `thread_url`, `garden_root_url`, `room_url`, `From`/`Into` `ThreadNav`. +- **Direction:** new code and refactored Maud should take **`&RouteContext`** (or owned where appropriate) instead of `&ThreadNav` when building links. Long term, **`item_href(&ItemId)`** should not parse strings — it should pattern-match `ItemId` and append tilde tail or `/-/…` external tail using the same rules as today’s `ThreadNav::garden_item_url`. + +### Axum / garden routes + +- **No** single catch-all route (explicit decision): keep the existing router layout in `server/src/lib.rs`. +- Room routes stay **`/r/:room_key/...`** with `room_key` fused; parsing via `slug_types::room_id_from_route_segment` / `room_route_segment` in `paths.rs`. + +## Phased execution (recommended order) + +### Phase 0 — Preconditions (quick) + +1. Read **`AGENTS.md`** (UI contract, durability matrix, `RpcCommand` vs `HtmlUiAction`). +2. Run **`cargo test --workspace`** and **`./scripts/clj-test.sh`** on clean `main` before large diffs; repeat after each phase. + +### Phase 1 — `ItemId` in `slug-types` (no server yet) + +1. Add **`ItemId`** (new file e.g. `types/src/item_id.rs` **or** inline at bottom of `paths.rs` — see **Module cycle** below). +2. Implement **`ItemId::parse`** using existing **`canonicalize_item`** + normalization; port **`CanonicalItemUrl`** methods to **`ItemId`** with tests ported from `paths.rs` `#[cfg(test)] mod tests`. +3. **`GardenItemUrl::from_stored(&ItemId, room_wire)`** (and thread helpers) — build absolute hrefs from structure, not from re-parsing a canonical string. +4. **`TildeHttpPathTail::to_item_id`** (rename from `to_canonical`) / **`tilde_http_path_to_item_id`**. +5. **`TildeOntologyPath::from_stored(&ItemId)`**. +6. Export **`ItemId`** from **`types/src/lib.rs`**; update **`server/src/path_types.rs`** re-exports. +7. **Delete `CanonicalItemUrl`** and fix all **in-crate** references in `types` only until `cargo test` passes for `slug-types`. + +**Module cycle trap:** `item_id.rs` must not `use crate::paths::{...}` if `paths.rs` also imports `ItemId` for `GardenItemUrl` in the same module. **Fix one of:** + +- **A)** Put `ItemId` **inside `paths.rs`** below `canonicalize_item` / helpers (simplest, large file), or +- **B)** Split **`canonicalize_item`** (+ dash host helpers + `finalize_external_identity_url`) into **`types/src/item_wire.rs`**, then `paths.rs` + `item_id.rs` both depend on `item_wire` only (cleaner, more files). + +### Phase 2 — Reducer + ranking (server core) + +1. **`server/src/reducer.rs`**: `ContentState` / `GroupState` / **`VoteData`** — replace **`CanonicalItemUrl`** with **`ItemId`** on all maps, sets, deques, vectors. +2. **`apply_vote`**: normalize `a`/`b` via **`ItemId::parse`** or **`ItemId`**-aware logic (remove string round-trip). +3. **`apply_ingest_to_content`**: **`dsl`** still yields strings for item titles in statements; normalize to **`ItemId`** at ingest boundary via **`ItemId::parse`** once per item. +4. **`server/src/ranking.rs`**, **`server/src/scope_rank.rs`**, **`server/src/api/write_actor.rs`** (including **`BTreeSet`** ordering), **`server/src/api/validate.rs`**, **`server/src/api/helpers.rs`** — propagate **`ItemId`**. +5. **`server/tests/basic.rs`** and any reducer tests constructing **`VoteData`** — use **`ItemId::parse(...).unwrap()`** or helpers. + +### Phase 3 — RPC + search + external resolver + +1. **`server/src/api/rpc.rs`**: rank/pair/matchup/search payloads; today many paths use **`GardenItemUrl::from_storage_str(item.as_str(), …)`** — switch to **`ItemId`** + **`GardenItemUrl::from_stored(&item_id, …)`** (or equivalent). +2. **`server/src/html/search.rs`**: scoring uses item path strings — derive from **`ItemId::display_path`** / **`to_wire_url`** only at the scoring boundary if needed. +3. **`server/src/external_resolver.rs`**: take **`&ItemId`** or **`ItemId::external_url()`** instead of **`&CanonicalItemUrl`**. + +### Phase 4 — HTML / Maud + +1. **`ThreadNav::garden_item_url`**: overload or replace with **`garden_item_href(&self, item: &ItemId)`** (no `CanonicalItemUrl::parse` inside). +2. **`RouteContext`**: extend **`item_href(&ItemId)`**; migrate call sites from **`ThreadNav`** to **`RouteContext`** where only link-building is needed (keep **`ThreadNav`** where scope / auth helpers need the full struct). +3. **`server/src/html/garden.rs`**, **`breadcrumb_path.rs`**, **`forum/*`**, **`editor.rs`**: replace **`CanonicalItemUrl`** with **`ItemId`**; breadcrumbs should walk **`ItemId::parent`** without string `rsplit`. +4. **`types` JSON types** (`RankRow`, etc.): decide whether **`GardenItemUrl`** stays string for JSON or becomes a structured field; keep **one** wire format for the public API. + +### Phase 5 — Cleanup + docs + +1. Remove dead **`canonical_path`** / **`breadcrumb_path`** string logic if fully superseded. +2. Update **`AGENTS.md`** if durability, `POST /ui`, or command surfaces change. +3. Delete or shrink **`plan.md`** when done. + +## File / symbol checklist (non-exhaustive — grep-driven) + +Run periodically: + +```bash +rg "CanonicalItemUrl" -g'*.rs' +rg "path_types::CanonicalItemUrl" -g'*.rs' +rg "tilde_http_path_to_canonical" -g'*.rs' +``` + +**High-touch files (from prior exploration):** + +| Area | Files | +|------|--------| +| Types | `types/src/paths.rs`, `types/src/lib.rs`, `types/src/url_normalize.rs`, (optional) `types/src/item_id.rs`, `types/src/item_wire.rs` | +| Server re-exports | `server/src/path_types.rs`, `server/src/canonical_path.rs` | +| Reducer / ingest | `server/src/reducer.rs`, `server/src/dsl.rs` (parse output types if changed) | +| Ranking | `server/src/ranking.rs`, `server/src/scope_rank.rs` | +| Writer / RPC | `server/src/api/write_actor.rs`, `server/src/api/rpc.rs`, `server/src/api/helpers.rs`, `server/src/api/validate.rs` | +| HTML | `server/src/html/garden.rs`, `server/src/html/breadcrumb_path.rs`, `server/src/html/forum/nav.rs`, `server/src/html/routing.rs`, `server/src/html/search.rs`, `server/src/html/editor.rs`, `server/src/html/forum/ingest.rs`, … | +| Tests | `server/tests/basic.rs`, `server/tests/integration.rs`, `types/src/paths.rs` tests, Clojure under `test/` if URLs/assertions mention canonical shapes | + +## Events / JSONL + +- **`Ingest`** events store **`raw` DSL** only — no change required for item identity inside the event. +- If any future event type stores item ids as strings, migrate to **structured `ItemId` serde** or accept string only at the event boundary with immediate parse into **`ItemId`** on `apply_event`. + +## `nav!` macro (`server/src/paths.rs`) + +- Macros use **`keypath($key)`** with **`.clone()`** — **`ItemId`** must be **`Clone`** (already for enums). Remove any reliance on **`Borrow`** for map keys. + +## Testing gate + +After each phase: + +```bash +cargo test --workspace +./scripts/clj-test.sh +``` + +## Risks / gotchas + +1. **`Ord` on `ItemId`**: must match prior **`CanonicalItemUrl`** / `String` ordering wherever **`BTreeSet`** is used (e.g. deterministic scope-rank snapshots in **`write_actor`**). +2. **External `ItemId`**: **`Url`** equality / hashing — normalization is already centralized in **`url_normalize`**; ensure **`ItemId::parse`** always inserts normalized **`Url`** into **`External`**. +3. **Fake parent URLs** in garden (e.g. **`https://.`** for external root ranking): find all **`parse("https://.")`** style hacks and express as **`ItemId`** or a dedicated sentinel. +4. **Serde**: tests and any RPC clients that snapshot JSON may need expectation updates if **`VoteData`** shape changes. + +## Optional follow-ups (not blocking `ItemId`) + +- More **domain normalizers** in **`url_normalize.rs`** (e.g. `music.youtube.com`, Spotify, etc.). +- **Room wire** vs **HTTP segment** helpers already in **`paths.rs`** (`ROOM_SHORT_ID_LEN`, `room_route_segment`, `room_id_from_route_segment`). + +--- + +**End state criteria:** `rg CanonicalItemUrl` returns nothing; reducer maps use **`ItemId`**; HTML link generation for items goes through **`RouteContext` + `ItemId`**; tests and Kaocha green. diff --git a/server/src/html/mod.rs b/server/src/html/mod.rs index c3dccd03dc6da2a2f6f6fa828657e33a951884b1..668403a35c415e6f091362b28c63ac294057fb6d 100644 --- a/server/src/html/mod.rs +++ b/server/src/html/mod.rs @@ -15,6 +15,7 @@ mod breadcrumb_path; mod editor; mod forum; mod garden; +pub mod routing; mod search; pub mod ui_action; use breadcrumb_path::{ExternalOntologyPath, OntologyPath}; @@ -35,6 +36,7 @@ pub use garden::{ external_garden_index, external_ontology_path, garden_index, ontology_path, room_external_garden_index, room_external_ontology_path, room_garden_index, room_ontology_path, }; +pub use routing::RouteContext; pub use search::{search_page, search_results_fragment}; pub use forum::user_profile_page; pub use ui_action::{parse_html_ui_from_form, HtmlUiAction, HtmlUiParseError, UI_RPC_FIELD}; diff --git a/server/src/html/routing.rs b/server/src/html/routing.rs new file mode 100644 index 0000000000000000000000000000000000000000..13e9ae0cbe32d454e34a7737edf8234d2a558502 --- /dev/null +++ b/server/src/html/routing.rs @@ -0,0 +1,71 @@ +//! Scoped browser paths for HTML. **RouteContext** is the intended single place to build `href`s +//! given public vs room scope (the original blueprint name); today it wraps [`ThreadNav`]. +//! +//! Prefer `RouteContext::item_href` / [`RouteContext::thread_url`] in new Maud over stitching +//! `/r/…` vs `/~` manually. Call sites can migrate incrementally from passing `&ThreadNav`. + +use crate::path_types::CanonicalItemUrl; + +use super::forum::ThreadNav; + +#[derive(Clone)] +pub struct RouteContext(ThreadNav); + +impl RouteContext { + #[inline] + pub fn public() -> Self { + Self(ThreadNav::public()) + } + + #[inline] + pub fn from_room_id(room_id: &str) -> Option { + ThreadNav::from_room_id(room_id).map(Self) + } + + #[inline] + pub fn thread_nav(&self) -> &ThreadNav { + &self.0 + } + + #[inline] + pub fn into_thread_nav(self) -> ThreadNav { + self.0 + } + + /// Relative path for a stored canonical item in this scope’s garden. + pub fn item_href(&self, item: &CanonicalItemUrl) -> String { + self.0.garden_item_url(item.as_str()) + } + + /// Same as [`Self::item_href`] but parses `item` first (raw DSL / user paste). + pub fn item_href_raw(&self, item: &str) -> String { + self.0.garden_item_url(item) + } + + #[inline] + pub fn thread_url(&self, tag: &str) -> String { + self.0.thread_url(tag) + } + + #[inline] + pub fn garden_root_url(&self) -> &str { + self.0.garden_root_url() + } + + #[inline] + pub fn room_url(&self) -> &str { + self.0.room_url() + } +} + +impl From for RouteContext { + fn from(nav: ThreadNav) -> Self { + Self(nav) + } +} + +impl From for ThreadNav { + fn from(ctx: RouteContext) -> Self { + ctx.0 + } +} Side B — contributor: tommy-mor Side B — commit message: [40b975bf] nice Side B — unified diff (full patch): diff --git a/Cargo.lock b/Cargo.lock index 266e876bb7ccbe788beb1d5bd53ad5b45ee5825b..2cea973082716e761ef6f5dd5886acc08ff9aac0 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -222,6 +222,12 @@ dependencies = [ "syn", ] +[[package]] +name = "dotenvy" +version = "0.15.7" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "1aaf95b3e5c8f23aa320147307562d361db0ae0d51242340f558153b4eb2439b" + [[package]] name = "encoding_rs" version = "0.8.35" @@ -1238,6 +1244,7 @@ version = "0.0.1" dependencies = [ "axum", "axum-extra", + "dotenvy", "maud", "reqwest", "serde", diff --git a/server/Cargo.toml b/server/Cargo.toml index 4677fedcb45292eebebe7e9cf6ce2f5738f18ddf..bd600138b613bd0f546bdec217a5334cdcb20aa5 100644 --- a/server/Cargo.toml +++ b/server/Cargo.toml @@ -17,6 +17,7 @@ tower-http = { version = "0.5", features = ["trace"] } tracing = "0.1" tracing-subscriber = { version = "0.3", features = ["env-filter"] } reqwest = { version = "0.12", features = ["json"] } +dotenvy = "0.15" [dev-dependencies] reqwest = { version = "0.12", features = ["json"] } diff --git a/server/src/api/ui_html.rs b/server/src/api/ui_html.rs index d2024bd4582bcc8482b461b2ba4fedbd8bff7c66..b33a84e8bb5e817b26592868d88090e6d664d950 100644 --- a/server/src/api/ui_html.rs +++ b/server/src/api/ui_html.rs @@ -6,7 +6,7 @@ use axum::{ use std::collections::HashMap; use crate::{ - html::{input_panel, js_string_literal, ranking_panel, JsBuilder}, + html::{entity_section, input_panel, js_string_literal, ranking_panel, JsBuilder}, parser::parse_reddit_url, path_types::ItemId, reddit::ensure_partial_tree, @@ -87,6 +87,20 @@ pub async fn post_ui_html( .into_response() } }, + HtmlUiAction::FetchEntity { item } => { + let id = parse_item_param(&item); + if id.is_root() { + return ui_js_warn("nothing to fetch for the root").into_response(); + } + state.queue_entity_fetch(id.clone()); + let tree = state.tree.read().await; + let empty = crate::reducer::NodeState::default(); + let node = tree.get(&id).unwrap_or(&empty); + let panel = entity_section(&id, node, true); + JsBuilder::new() + .morph_selector("#entity-section", panel) + .into_response() + }, } } diff --git a/server/src/events.rs b/server/src/events.rs index ed5be6b13b9d46e838831d6ce0f96f569b401730..07ce24b5e56cf72b0b442c3c3241efbf6c3b006a 100644 --- a/server/src/events.rs +++ b/server/src/events.rs @@ -1,4 +1,5 @@ use serde::{Deserialize, Serialize}; +use serde_json::Value; #[derive(Debug, Clone, Serialize, Deserialize)] #[serde(tag = "type", rename_all = "snake_case")] @@ -18,4 +19,10 @@ pub enum Event { }, /// Register a node path in the fractal tree (no external fetch). NodeEnsured { id: String }, + /// Full upstream API payload for a node (domain-specific view derived at replay/render time). + EntityImported { + id: String, + ts: i64, + payload: Value, + }, } diff --git a/server/src/html/mod.rs b/server/src/html/mod.rs index df6505021d9f446c2b453e20e3eb3cf696a111f9..caf1309c8d93b47104499c57f9cc35ee7631fbb9 100644 --- a/server/src/html/mod.rs +++ b/server/src/html/mod.rs @@ -10,6 +10,7 @@ use crate::{ form_template::template_json_compact, path_types::ItemId, ranking::{top_bottom, RankedItem}, + reddit::is_fetchable, reducer::{GroupState, NodeState}, state::AppState, ui_action::UI_RPC_FIELD, @@ -151,7 +152,7 @@ pub fn breadcrumb_path(item: &ItemId) -> Markup { fn entity_panel(node: &NodeState) -> Markup { html! { @if let Some(data) = &node.data { - section id="entity-panel" class="demo-panel entity-card" { + div id="entity-panel" class="entity-card" { h2 { (data.title) } @if let Some(author) = &data.author { p class="muted small" { "by " (author) } @@ -164,6 +165,42 @@ fn entity_panel(node: &NodeState) -> Markup { } } +/// Reddit/API import control — only shown on fetchable pages; never auto-fires. +pub fn fetch_entity_panel(item: &ItemId, has_data: bool, fetching: bool) -> Markup { + if !is_fetchable(item) { + return html! {}; + } + let label = if fetching { + "Fetching…" + } else if has_data { + "Fetch more" + } else { + "Fetch from Reddit" + }; + let rpc = template_json_compact(&serde_json::json!({ + "action": "fetch_entity", + "item": item.as_str(), + })) + .expect("fetch_entity rpc template"); + html! { + form method="post" action="/ui" id="fetch-entity-form" class="fetch-entity-form" { + input type="hidden" name=(UI_RPC_FIELD) value=(rpc); + button type="submit" class="btn-secondary" disabled=(fetching) { (label) } + } + } +} + +/// Entity card + explicit fetch control (morphed as `#entity-section`). +pub fn entity_section(item: &ItemId, node: &NodeState, fetching: bool) -> Markup { + let has_data = node.data.is_some(); + html! { + section id="entity-section" class="demo-panel" { + (entity_panel(node)) + (fetch_entity_panel(item, has_data, fetching)) + } + } +} + fn rank_list(label: &str, items: &[RankedItem], start_rank: usize) -> Markup { html! { @if !items.is_empty() { @@ -260,7 +297,7 @@ async fn item_page(state: AppState, uri: Uri, item: ItemId) -> Markup { h1 { "sorter" } (input_panel("", None)) (breadcrumb_path(&item)) - (entity_panel(node)) + (entity_section(&item, node, false)) (ranking_panel(&item, group)) }; layout("sorter2", body, views) @@ -272,16 +309,5 @@ pub async fn home(State(state): State, uri: Uri) -> impl IntoResponse pub async fn browse(State(state): State, uri: Uri) -> impl IntoResponse { let item = ItemId::from_browse_uri(uri.path()).unwrap_or(ItemId::root()); - if item.as_str().starts_with("reddit.com") { - let needs_fetch = { - let tree = state.tree.read().await; - tree.get(&item) - .map(|n| n.data.is_none()) - .unwrap_or(true) - }; - if needs_fetch { - state.reddit.request_fetch(item.clone()); - } - } item_page(state, uri, item).await } diff --git a/server/src/main.rs b/server/src/main.rs index c22ec6c9f5358e5ec99fb83210dc351938505a93..1f0cddc39302b35b0cd6a6219f44c9d59202facf 100644 --- a/server/src/main.rs +++ b/server/src/main.rs @@ -2,6 +2,10 @@ use sorter2_server::state::AppConfig; #[tokio::main] async fn main() -> Result<(), Box> { + if std::env::var("SORTER2_SKIP_DOTENV").is_err() { + let _ = dotenvy::dotenv(); + } + tracing_subscriber::fmt() .with_env_filter( tracing_subscriber::EnvFilter::try_from_default_env() diff --git a/server/src/reddit.rs b/server/src/reddit.rs index 90053ad03b1d7c8e94f325dd4ee64c2b4f7da900..ff0f01e57b18af878eb5be3efc47204a7673589d 100644 --- a/server/src/reddit.rs +++ b/server/src/reddit.rs @@ -6,11 +6,15 @@ use std::time::{Duration, Instant}; use reqwest::{header, Client, StatusCode}; use serde::Deserialize; +use serde_json::Value; use tokio::sync::{mpsc, RwLock}; use crate::{ + event_log::EventLog, + events::Event, + html::now_ms, path_types::ItemId, - reducer::{EntityData, GlobalTree}, + reducer::GlobalTree, }; /// Bootstrap blank nodes along a URL path so breadcrumbs and voting work before fetch. @@ -20,6 +24,8 @@ pub fn ensure_partial_tree(tree: &mut GlobalTree, id: &ItemId) { pub struct RedditCommand { pub id: ItemId, + /// User-initiated fetch bypasses the in-memory "recently fetched" cache. + pub force: bool, } #[derive(Clone)] @@ -33,19 +39,31 @@ struct RedditCredentials { client_secret: String, } +#[derive(Clone)] +pub struct RedditApiConfig { + pub api_base: String, + pub oauth_base: String, + pub user_agent: String, + creds: Option, +} + struct OAuthToken { access_token: String, expires_at: Instant, } impl RedditBroker { - pub fn spawn(tree: Arc>, user_agent: &str) -> Self { + pub fn spawn( + tree: Arc>, + event_log: Arc, + config: RedditApiConfig, + ) -> Self { let (tx, rx) = mpsc::channel(100); let mut headers = header::HeaderMap::new(); headers.insert( header::USER_AGENT, - header::HeaderValue::from_str(user_agent).expect("valid user agent"), + header::HeaderValue::from_str(&config.user_agent).expect("valid user agent"), ); let client = Client::builder() @@ -54,22 +72,38 @@ impl RedditBroker { .build() .expect("reqwest client"); - let creds = RedditCredentials::from_env(); - tokio::spawn(reddit_worker(rx, tree, client, creds)); + tokio::spawn(reddit_worker(rx, tree, event_log, client, config)); Self { tx } } - /// Fire-and-forget: queue a fetch; worker updates the tree when done. - pub fn request_fetch(&self, id: ItemId) { - let _ = self.tx.try_send(RedditCommand { id }); + /// Queue a fetch; drops when the channel is full (backpressure). + pub fn request_fetch(&self, id: ItemId, force: bool) { + let _ = self.tx.try_send(RedditCommand { id, force }); + } +} + +impl RedditApiConfig { + pub fn from_env() -> Self { + Self { + api_base: reddit_api_base(), + oauth_base: reddit_oauth_base(), + user_agent: default_user_agent(), + creds: RedditCredentials::from_env(), + } } } impl RedditCredentials { + /// Reddit's OAuth docs call these "client id" and "client secret"; the app + /// registration UI often labels them "app id" / "app secret" — same values. fn from_env() -> Option { - let client_id = std::env::var("REDDIT_CLIENT_ID").ok()?; - let client_secret = std::env::var("REDDIT_CLIENT_SECRET").ok()?; + let client_id = std::env::var("REDDIT_CLIENT_ID") + .or_else(|_| std::env::var("REDDIT_APP_ID")) + .ok()?; + let client_secret = std::env::var("REDDIT_CLIENT_SECRET") + .or_else(|_| std::env::var("REDDIT_APP_SECRET")) + .ok()?; if client_id.is_empty() || client_secret.is_empty() { return None; } @@ -80,29 +114,63 @@ impl RedditCredentials { } } +pub fn reddit_api_base() -> String { + std::env::var("REDDIT_API_BASE").unwrap_or_else(|_| "https://www.reddit.com".into()) +} + +pub fn reddit_oauth_base() -> String { + std::env::var("REDDIT_OAUTH_BASE").unwrap_or_else(|_| "https://www.reddit.com".into()) +} + pub fn default_user_agent() -> String { std::env::var("REDDIT_USER_AGENT").unwrap_or_else(|_| { "web:sorter2.social:v0.0.1 (by /u/sorter2)".to_string() }) } +/// True when this node can be loaded from the Reddit JSON API. +pub fn is_fetchable(id: &ItemId) -> bool { + !map_item_to_reddit_api(id, "https://example.com").is_empty() +} + +/// Derive UI-facing fields from a stored payload (Reddit-specific when under reddit.com). +pub fn entity_view_from_payload(id: &ItemId, payload: &Value) -> Option { + if id.as_str().starts_with("reddit.com") { + return parse_reddit_view(id, payload); + } + None +} + +/// Apply a full API payload to the in-memory tree (view derived for known domains). +pub fn apply_entity_import(tree: &mut GlobalTree, id: &ItemId, payload: Value) { + let view = entity_view_from_payload(id, &payload); + tree.apply_entity_raw(id, payload, view); +} + async fn reddit_worker( mut rx: mpsc::Receiver, tree: Arc>, + event_log: Arc, client: Client, - creds: Option, + config: RedditApiConfig, ) { let mut in_flight = HashSet::new(); let mut recently_fetched: HashMap = HashMap::new(); let mut current_delay = Duration::from_secs(1); let mut oauth: Option = None; let cache_ttl = Duration::from_secs(300); + let creds = config.creds.clone(); + let api_base = config.api_base.clone(); + let oauth_base = config.oauth_base.clone(); while let Some(cmd) = rx.recv().await { let now = Instant::now(); recently_fetched.retain(|_, t| now.duration_since(*t) < cache_ttl); - if in_flight.contains(&cmd.id) || recently_fetched.contains_key(&cmd.id) { + if in_flight.contains(&cmd.id) { + continue; + } + if !cmd.force && recently_fetched.contains_key(&cmd.id) { continue; } @@ -112,18 +180,26 @@ async fn reddit_worker( tokio::time::sleep(current_delay).await; if let Some(c) = &creds { - oauth = ensure_oauth_token(&client, c, oauth.take()).await; + oauth = ensure_oauth_token(&client, &oauth_base, c, oauth.take()).await; } let token = oauth.as_ref().map(|t| t.access_token.as_str()); - let use_oauth = token.is_some(); - match do_fetch(&client, &fetch_id, use_oauth, token).await { - Ok(FetchOutcome::Entity(data)) => { - let mut w = tree.write().await; - w.set_entity_data(&fetch_id, data); - recently_fetched.insert(fetch_id.clone(), Instant::now()); - current_delay = Duration::from_millis(600); + match do_fetch(&client, &api_base, &fetch_id, token).await { + Ok(FetchOutcome::Payload(payload)) => { + let ts = now_ms(); + let event = Event::EntityImported { + id: fetch_id.as_str().to_string(), + ts, + payload: payload.clone(), + }; + if let Err(e) = event_log.append(&event).await { + tracing::warn!("event log append failed for {}: {}", fetch_id, e); + } else { + apply_entity_import(&mut *tree.write().await, &fetch_id, payload); + recently_fetched.insert(fetch_id.clone(), Instant::now()); + current_delay = Duration::from_millis(600); + } } Ok(FetchOutcome::NotFound) => { recently_fetched.insert(fetch_id.clone(), Instant::now()); @@ -149,13 +225,14 @@ async fn reddit_worker( } enum FetchOutcome { - Entity(EntityData), + Payload(Value), NotFound, RateLimited { reset_secs: u64 }, } async fn ensure_oauth_token( client: &Client, + oauth_base: &str, creds: &RedditCredentials, existing: Option, ) -> Option { @@ -165,8 +242,13 @@ async fn ensure_oauth_token( } } + let url = format!( + "{}/api/v1/access_token", + oauth_base.trim_end_matches('/') + ); + let resp = client - .post("https://www.reddit.com/api/v1/access_token") + .post(&url) .basic_auth(&creds.client_id, Some(&creds.client_secret)) .form(&[("grant_type", "client_credentials")]) .send() @@ -207,11 +289,11 @@ async fn ensure_oauth_token( async fn do_fetch( client: &Client, + api_base: &str, id: &ItemId, - use_oauth: bool, bearer: Option<&str>, ) -> Result { - let url = map_item_to_reddit_api(id, use_oauth); + let url = map_item_to_reddit_api(id, api_base); if url.is_empty() { return Ok(FetchOutcome::NotFound); } @@ -241,10 +323,8 @@ async fn do_fetch( return Ok(FetchOutcome::RateLimited { reset_secs: reset }); } - let bytes = resp.bytes().await.map_err(|e| e.to_string())?; - Ok(parse_reddit_json(id, &bytes) - .map(FetchOutcome::Entity) - .unwrap_or(FetchOutcome::NotFound)) + let payload: Value = resp.json().await.map_err(|e| e.to_string())?; + Ok(FetchOutcome::Payload(payload)) } fn rate_limit_remaining(resp: &reqwest::Response) -> Option { @@ -264,18 +344,14 @@ fn rate_limit_reset_secs(resp: &reqwest::Response) -> u64 { .unwrap_or(5) } -/// Map canonical item id to Reddit JSON API URL. -pub fn map_item_to_reddit_api(id: &ItemId, oauth: bool) -> String { +/// Map canonical item id to a Reddit JSON API URL under `api_base`. +pub fn map_item_to_reddit_api(id: &ItemId, api_base: &str) -> String { let path = id.as_str(); if !path.starts_with("reddit.com/") && path != "reddit.com" { return String::new(); } - let base = if oauth { - "https://oauth.reddit.com" - } else { - "https://www.reddit.com" - }; + let base = api_base.trim_end_matches('/'); let segments: Vec<&str> = path.split('/').collect(); @@ -293,18 +369,17 @@ pub fn map_item_to_reddit_api(id: &ItemId, oauth: bool) -> String { String::new() } -fn parse_reddit_json(id: &ItemId, bytes: &[u8]) -> Option { - let v: serde_json::Value = serde_json::from_slice(bytes).ok()?; +fn parse_reddit_view(id: &ItemId, v: &Value) -> Option { let segments: Vec<&str> = id.as_str().split('/').collect(); if segments.iter().any(|&p| p == "comments") { - parse_post_listing(&v) + parse_post_listing(v) } else { - parse_subreddit_about(&v) + parse_subreddit_about(v) } } -fn parse_subreddit_about(v: &serde_json::Value) -> Option { +fn parse_subreddit_about(v: &Value) -> Option { let data = v.get("data")?; let title = data .get("title") @@ -323,7 +398,7 @@ fn parse_subreddit_about(v: &serde_json::Value) -> Option { .filter(|s| !s.is_empty()) .map(|s| s.to_string()); - Some(EntityData { + Some(crate::reducer::EntityData { title, author: None, body_html, @@ -331,10 +406,9 @@ fn parse_subreddit_about(v: &serde_json::Value) -> Option { }) } -fn parse_post_listing(v: &serde_json::Value) -> Option { +fn parse_post_listing(v: &Value) -> Option { let listing = v.as_array()?.first()?; - let child = listing - .pointer("/data/children/0/data")?; + let child = listing.pointer("/data/children/0/data")?; let title = child.get("title")?.as_str()?.to_string(); let author = child .get("author") @@ -352,7 +426,7 @@ fn parse_post_listing(v: &serde_json::Value) -> Option { .filter(|s| s.starts_with("http")) .map(|s| s.to_string()); - Some(EntityData { + Some(crate::reducer::EntityData { title, author, body_html, @@ -368,48 +442,51 @@ mod tests { fn map_subreddit_about_url() { let id = ItemId::parse("reddit.com/r/rust").unwrap(); assert_eq!( - map_item_to_reddit_api(&id, false), + map_item_to_reddit_api(&id, "https://www.reddit.com"), "https://www.reddit.com/r/rust/about.json?raw_json=1" ); assert_eq!( - map_item_to_reddit_api(&id, true), - "https://oauth.reddit.com/r/rust/about.json?raw_json=1" + map_item_to_reddit_api(&id, "http://127.0.0.1:9999"), + "http://127.0.0.1:9999/r/rust/about.json?raw_json=1" ); } #[test] fn map_post_url() { - let id = - ItemId::parse("reddit.com/r/amitheasshole/comments/1trnvdl").unwrap(); + let id = ItemId::parse("reddit.com/r/amitheasshole/comments/1trnvdl").unwrap(); assert_eq!( - map_item_to_reddit_api(&id, false), + map_item_to_reddit_api(&id, "https://www.reddit.com"), "https://www.reddit.com/r/amitheasshole/comments/1trnvdl.json?raw_json=1" ); } #[test] - fn map_non_reddit_empty() { - let id = ItemId::opaque("example.com/foo"); - assert!(map_item_to_reddit_api(&id, false).is_empty()); + fn is_fetchable_reddit_sub() { + let id = ItemId::parse("reddit.com/r/rust").unwrap(); + assert!(is_fetchable(&id)); + assert!(!is_fetchable(&ItemId::opaque("example.com/x"))); } #[test] fn parse_subreddit_fixture() { - let json = r#"{"kind":"t5","data":{"title":"Rust","display_name":"rust","public_description":"systems"}}"#; - let entity = parse_reddit_json( + let json = include_str!("../../test/fixtures/reddit/r_rust_about.json"); + let v: Value = serde_json::from_str(json).unwrap(); + let entity = entity_view_from_payload( &ItemId::parse("reddit.com/r/rust").unwrap(), - json.as_bytes(), + &v, ) .unwrap(); - assert_eq!(entity.title, "Rust"); + assert_eq!(entity.title, "The Rust Programming Language"); + assert!(entity.body_html.as_ref().is_some_and(|b| b.contains("Rust"))); } #[test] fn parse_post_fixture() { let json = r#"[{"kind":"Listing","data":{"children":[{"kind":"t3","data":{"title":"AITA","author":"op","selftext_html":"<p>hi</p>","thumbnail":"https://b.thumbs.redditmedia.com/x.jpg"}}]}}]"#; - let entity = parse_reddit_json( + let v: Value = serde_json::from_str(json).unwrap(); + let entity = entity_view_from_payload( &ItemId::parse("reddit.com/r/x/comments/abc").unwrap(), - json.as_bytes(), + &v, ) .unwrap(); assert_eq!(entity.title, "AITA"); diff --git a/server/src/reducer.rs b/server/src/reducer.rs index 077f700bf00ddefe18ffd004bb5288bdc7c4adaf..60db81a562e3590775012573225a76ba82a14563 100644 --- a/server/src/reducer.rs +++ b/server/src/reducer.rs @@ -1,6 +1,7 @@ use std::collections::{HashMap, HashSet, VecDeque}; use serde::{Deserialize, Serialize}; +use serde_json::Value; use crate::path_types::ItemId; @@ -128,6 +129,9 @@ pub struct EntityData { #[derive(Debug, Clone, Default)] pub struct NodeState { pub id: ItemId, + /// Full imported API JSON (persisted in the event log). + pub entity_raw: Option, + /// Domain-specific view derived from `entity_raw` (e.g. Reddit title/author). pub data: Option, pub children: HashSet, pub local_ranking: GroupState, @@ -197,10 +201,11 @@ impl GlobalTree { } } - pub fn set_entity_data(&mut self, id: &ItemId, data: EntityData) { + pub fn apply_entity_raw(&mut self, id: &ItemId, payload: Value, view: Option) { self.ensure_path(id); if let Some(node) = self.nodes.get_mut(id) { - node.data = Some(data); + node.entity_raw = Some(payload); + node.data = view; } } } diff --git a/server/src/state.rs b/server/src/state.rs index 8c03aa60c15aee803a534439400b69935b1a3d84..d71d1079a8f486cdab15795384aef0b81b32544d 100644 --- a/server/src/state.rs +++ b/server/src/state.rs @@ -7,7 +7,7 @@ use crate::{ events::Event, journal::JournalClient, path_types::ItemId, - reddit::{default_user_agent, RedditBroker}, + reddit::{apply_entity_import, RedditApiConfig, RedditBroker}, reducer::{GlobalTree, VoteData}, views::ViewStore, }; @@ -108,13 +108,18 @@ impl AppState { tree.ensure_path(&parsed); } } + Event::EntityImported { id, payload, .. } => { + if let Some(parsed) = ItemId::parse(&id).or_else(|| ItemId::from_url(&id)) { + apply_entity_import(&mut tree, &parsed, payload); + } + } } } } let tree = Arc::new(RwLock::new(tree)); let journal = JournalClient::spawn(tree.clone(), event_log.clone()); - let reddit = RedditBroker::spawn(tree.clone(), &default_user_agent()); + let reddit = RedditBroker::spawn(tree.clone(), event_log.clone(), RedditApiConfig::from_env()); Self { cfg: Arc::new(cfg), @@ -135,10 +140,14 @@ impl AppState { let mut w = self.tree.write().await; w.ensure_path(id); } - self.reddit.request_fetch(id.clone()); Ok(()) } + /// User-initiated Reddit/API import (via "Fetch more" — never on paste or navigate). + pub fn queue_entity_fetch(&self, id: ItemId) { + self.reddit.request_fetch(id, true); + } + pub async fn record_vote( &self, parent: &ItemId, @@ -169,6 +178,44 @@ impl AppState { #[cfg(test)] mod tests { use super::{normalize_scope, parse_item_param}; + use crate::{ + event_log::EventLog, + events::Event, + path_types::ItemId, + reddit::apply_entity_import, + reducer::GlobalTree, + }; + use serde_json::json; + + #[tokio::test] + async fn replay_entity_imported_restores_view() { + let tmp = tempfile::tempdir().unwrap(); + let log_path = tmp.path().join("events.jsonl"); + let log = EventLog::new(log_path.to_string_lossy().into_owned()); + let payload = json!({"kind":"t5","data":{"title":"Rust","display_name":"rust"}}); + log.append(&Event::EntityImported { + id: "reddit.com/r/rust".into(), + ts: 1, + payload: payload.clone(), + }) + .await + .unwrap(); + + let mut tree = GlobalTree::new(); + let (events, _) = log.load_all().await.unwrap(); + for ev in events { + if let Event::EntityImported { id, payload, .. } = ev { + let parsed = ItemId::parse(&id).unwrap(); + apply_entity_import(&mut tree, &parsed, payload); + } + } + let node = tree.get(&ItemId::parse("reddit.com/r/rust").unwrap()).unwrap(); + assert_eq!(node.data.as_ref().unwrap().title, "Rust"); + assert_eq!( + node.entity_raw.as_ref().unwrap()["data"]["display_name"], + "rust" + ); + } #[test] fn normalize_scope_strips_prefix_and_lowercases() { diff --git a/server/src/ui_action.rs b/server/src/ui_action.rs index d798874d1d94c0dfee59ec1ff703f9ee6ef432c0..3d6a49a2a3fb950752827efe1f5a049308f09baf 100644 --- a/server/src/ui_action.rs +++ b/server/src/ui_action.rs @@ -26,6 +26,10 @@ pub enum HtmlUiAction { ParseQuery { query: String, }, + /// Fetch upstream entity data for the current page (explicit user action only). + FetchEntity { + item: String, + }, } #[derive(Debug, Error)] diff --git a/server/static/sorter.css b/server/static/sorter.css index 36f68d415cea33c8562d5d02c0c9d4c5b225a782..1f9f5158414902d43a380dd899232c70d8cf5d63 100644 --- a/server/static/sorter.css +++ b/server/static/sorter.css @@ -52,6 +52,25 @@ body { filter: brightness(1.1); } +.btn-secondary { + background: transparent; + color: var(--accent); + border: 1px solid var(--border); + padding: 0.4rem 0.85rem; + font-size: 0.9rem; + cursor: pointer; + margin-top: 0.75rem; +} + +.btn-secondary:disabled { + opacity: 0.6; + cursor: wait; +} + +.fetch-entity-form { + margin-top: 0.5rem; +} + code { font-size: 0.85em; background: var(--bg); diff --git a/test/fixtures/reddit/r_rust_about.json b/test/fixtures/reddit/r_rust_about.json new file mode 100644 index 0000000000000000000000000000000000000000..bb7fd74ab00d63356d9b672a775996ed155b984d --- /dev/null +++ b/test/fixtures/reddit/r_rust_about.json @@ -0,0 +1,21 @@ +{ + "kind": "t5", + "data": { + "user_flair_background_color": null, + "submit_text_html": null, + "banner_img": "", + "user_is_banned": null, + "wiki_enabled": true, + "show_media": true, + "id": "2qh0i", + "display_name": "rust", + "title": "The Rust Programming Language", + "public_description": "A place for all things related to the Rust programming language.", + "public_description_html": "<!-- SC_OFF --><div class=\"md\"><p>A place for all things related to the Rust programming language.</p>\n</div><!-- SC_ON -->", + "icon_img": "https://styles.redditmedia.com/t5_2qh0i/styles/communityIcon_kfqpknb2j6j51.png", + "community_icon": "https://styles.redditmedia.com/t5_2qh0i/styles/communityIcon_kfqpknb2j6j51.png", + "subscribers": 350000, + "active_user_count": 1200, + "over18": false + } +} diff --git a/test/reddit_import.clj b/test/reddit_import.clj new file mode 100644 index 0000000000000000000000000000000000000000..54edaedeef08718c8f184d6aa468d8e6415068c7 --- /dev/null +++ b/test/reddit_import.clj @@ -0,0 +1,116 @@ +(ns test.reddit-import + (:require [babashka.process :as process] + [clojure.java.io :as io] + [clojure.string :as str] + [clojure.test :refer [deftest is testing]]) + (:import [com.sun.net.httpserver HttpServer HttpHandler HttpExchange] + [java.net InetSocketAddress])) + +(defn- repo-root [] + (.getCanonicalPath (io/file (System/getProperty "user.dir")))) + +(defn- pick-port [] + (with-open [s (java.net.ServerSocket. 0)] + (.getLocalPort s))) + +(defn- start-mock-reddit [port fixtures-dir] + (let [fixture (io/file fixtures-dir "r_rust_about.json") + body (.getBytes (slurp fixture) "UTF-8") + server (HttpServer/create (InetSocketAddress. "127.0.0.1" port) 0) + handler + (proxy [HttpHandler] [] + (handle [^HttpExchange exchange] + (.sendResponseHeaders exchange 200 (alength body)) + (let [out (.getResponseBody exchange)] + (.write out body) + (.close out))))] + (.createContext server "/" handler) + (.setExecutor server nil) + (.start server) + (fn stop [] + (.stop server 0)))) + +(defn- wait-health [base-url ms] + (let [deadline (+ (System/currentTimeMillis) ms) + url (str base-url "/healthz")] + (loop [] + (let [resp (try + (process/shell {:out :string :err :string} + "curl" "-sf" url) + (catch Exception _ nil))] + (if (and resp (zero? (:exit resp)) (= "ok" (str/trim (:out resp "")))) + true + (if (< (System/currentTimeMillis) deadline) + (do (Thread/sleep 200) (recur)) + false)))))) + +(defn- curl-post-ui [base rpc-json] + (process/shell {:out :string :err :string} + "curl" "-sf" "-X" "POST" (str base "/ui") + "--data-urlencode" (str "__rpc__=" rpc-json))) + +(defn- wait-event-log [path ms] + (let [deadline (+ (System/currentTimeMillis) ms)] + (loop [] + (if (.exists (io/file path)) + true + (if (< (System/currentTimeMillis) deadline) + (do (Thread/sleep 200) (recur)) + false))))) + +(deftest reddit-fetch-via-mock-api + (testing "Fetch more queues import; event log stores full payload; page shows title" + (let [root (repo-root) + fixtures (str root "/test/fixtures/reddit") + data-dir (.getAbsolutePath + (doto (io/file (System/getProperty "java.io.tmpdir") + (str "sorter2-reddit-" (System/currentTimeMillis))) + (.mkdirs))) + reddit-port (pick-port) + app-port (pick-port) + reddit-base (str "http://127.0.0.1:" reddit-port) + app-base (str "http://127.0.0.1:" app-port) + bin (str root "/target/release/sorter2-server") + stop-mock (start-mock-reddit reddit-port fixtures)] + (try + (is (zero? (:exit (process/shell {:dir root} + "cargo" "build" "--release" "--package" "sorter2-server"))) + "release build succeeds") + (let [proc (process/process {:dir root + :env (into (into {} (System/getenv)) + {"SORTER2_SKIP_DOTENV" "1" + "SORTER2_DATA_DIR" data-dir + "SORTER2_EVENT_LOG" (str data-dir "/events.jsonl") + "PORT" (str app-port) + "REDDIT_API_BASE" reddit-base + "REDDIT_OAUTH_BASE" reddit-base + "REDDIT_CLIENT_ID" "" + "REDDIT_CLIENT_SECRET" "" + "REDDIT_APP_ID" "" + "REDDIT_APP_SECRET" ""}) + :out :string + :err :string} + bin)] + (try + (is (wait-health app-base 20000) "app healthz") + (let [browse-url (str app-base "/~/https://reddit.com/r/rust") + before (:out (process/shell {:out :string :err :string} + "curl" "-sf" browse-url))] + (is (str/includes? before "Fetch from Reddit")) + (is (not (str/includes? before "The Rust Programming Language"))) + (let [rpc "{\"action\":\"fetch_entity\",\"item\":\"reddit.com/r/rust\"}" + post (curl-post-ui app-base rpc) + log-path (str data-dir "/events.jsonl")] + (is (zero? (:exit post)) "fetch_entity POST succeeds") + (is (wait-event-log log-path 10000) "event log written") + (let [after (:out (process/shell {:out :string :err :string} + "curl" "-sf" browse-url)) + log (slurp (io/file log-path))] + (is (str/includes? after "The Rust Programming Language")) + (is (str/includes? log "\"type\":\"entity_imported\"")) + (is (str/includes? log "\"subscribers\":350000")) + (is (str/includes? log "\"display_name\":\"rust\""))))) + (finally + (process/destroy proc)))) + (finally + (stop-mock))))))