Side B implements a substantial, functional feature (a real Reddit fetch worker with OAuth, rate-limit backoff, dedup, JSON parsing, and unit tests) that provides real lasting functionality wired into the app state and HTML handler. Side A is a useful but small config simplification (collapsing two kaocha suites into one glob pattern), which is a low-risk convenience change with much less substance and no new tested behavior.
constitution · epochs · watch · epoch 3
c_28dfd8015c8a (tommy-mor) vs c_c124c217f89c (tommy-mor)
download prompt · raw event · cmp_4a739789a93b38
council reasoning
B implements a full Reddit fetch pipeline (broker, OAuth, rate-limit/backoff, JSON parsing, tree integration, and tests), turning a stub into working product behavior. A only rewrites tests.edn to a single auto-discovery suite—useful hygiene, but a small config change with far less lasting system impact.
Side B implements substantial new functionality: a background Reddit import system with request queuing, OAuth support, rate-limit handling, JSON parsing, integration into application state and request flow, plus unit tests for URL mapping and parsing. Side A only changes Kaocha configuration from manually enumerated test namespaces to a single auto-discovered suite, which improves maintainability but does not add comparable runtime capability.
sides
A — c_28dfd8015c8a (tommy-mor)
message
[0bebe819] Unify kaocha test discovery to one auto-discovered suite. Replace the two hand-enumerated suites (:http-integration, :browser) with a single :all suite that picks up every test.* namespace under test/. New test files now run automatically without touching tests.edn — needed for the test.ranking namespace added alongside the #146 fix, and for any future tests. Both ./TEST.sh and CI (.github/workflows/ci.yml) invoke `clojure -M:kaocha`, so both pick up the new suite without changes. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
diff preview
diff --git a/tests.edn b/tests.edn
index 5e2b0433a66c51806d08ef38d691259cfccd1fe0..e51ffcc81eeae6f41c85d1bfc5ba32f592cc5818 100644
--- a/tests.edn
+++ b/tests.edn
@@ -1,30 +1,14 @@
#kaocha/v1
-{:tests
- [{:id :http-integration
- :test-paths ["test"]
- :source-paths ["."]
- :ns-patterns ["^test\\.integration$"
- "^test\\.auth$"
- "^test\\.grants$"
- "^test\\.invites$"
- "^test\\.room-list$"]
- :kaocha.filter/skip-meta [:skip]
- :parallel? false}
- {:id :browser
- :test-paths ["test"]
- :source-paths ["."]
- :ns-patterns ["^test\\.browser-sse$"
- "^test\\.browser-ui-morph$"
- "^test\\.browser-post-redact$"
- "^test\\.browser-room-delete$"
- "^test\\.browser-public-garden$"
- "^test\\.browser-redact-thread-index$"
- "^test\\.browser-garden-pin$"
- "^test\\.browser-vote-compare$"
- "^test\\.browser-github-resolver$"]
- :kaocha.filter/skip-meta [:skip]
- :parallel? false}]
- :plugins [:kaocha.plugin/junit-xml]
- :kaocha.plugin.junit-xml/target-file "target/kaocha-junit.xml"
- :kaocha.plugin.junit-xml/add-location-metadata? true
- :reporter kaocha.report.progress/report}
+ {:tests
+ [{:id :all
+ :test-paths ["test"]
+ :source-paths ["."]
+ ;; Pick up every test.* namespace under test/. New files don't need to be
+ ;; enumerated here — drop them in test/ with `(ns test.foo …)` and they run.
+ :ns-patterns ["^test\\..+"]
+ :kaocha.filter/skip-meta [:skip]
+ :parallel? false}]
+ :plugins [:kaocha.plugin/junit-xml]
+ :kaocha.plugin.junit-xml/target-file "target/kaocha-junit.xml"
+ :kaocha.plugin.junit-xml/add-location-metadata? true
+ :reporter kaocha.report.progress/report}
B — c_c124c217f89c (tommy-mor)
message
[8d8230d1] reddit
diff preview
diff --git a/.gitignore b/.gitignore
index 4c7073f9fac0c30fd2050d79a60ef447af58ebeb..ada462e900d24a3a6d08165d158c80f79c35a5a5 100644
--- a/.gitignore
+++ b/.gitignore
@@ -7,3 +7,4 @@
data/
repomix-output.xml
dev-data/
+.env
diff --git a/server/Cargo.toml b/server/Cargo.toml
index 7906a8547d56b8e6a48ef59c37aa82a8510fdee9..4677fedcb45292eebebe7e9cf6ce2f5738f18ddf 100644
--- a/server/Cargo.toml
+++ b/server/Cargo.toml
@@ -16,6 +16,7 @@ tower = "0.5"
tower-http = { version = "0.5", features = ["trace"] }
tracing = "0.1"
tracing-subscriber = { version = "0.3", features = ["env-filter"] }
+reqwest = { version = "0.12", features = ["json"] }
[dev-dependencies]
reqwest = { version = "0.12", features = ["json"] }
diff --git a/server/src/html/mod.rs b/server/src/html/mod.rs
index 2864407ed6e8ec284a1dc663acf1805534566b62..df6505021d9f446c2b453e20e3eb3cf696a111f9 100644
--- a/server/src/html/mod.rs
+++ b/server/src/html/mod.rs
@@ -272,5 +272,16 @@ pub async fn home(State(state): State<AppState>, uri: Uri) -> impl IntoResponse
pub async fn browse(State(state): State<AppState>, uri: Uri) -> impl IntoResponse {
let item = ItemId::from_browse_uri(uri.path()).unwrap_or(ItemId::root());
+ if item.as_str().starts_with("reddit.com") {
+ let needs_fetch = {
+ let tree = state.tree.read().await;
+ tree.get(&item)
+ .map(|n| n.data.is_none())
+ .unwrap_or(true)
+ };
+ if needs_fetch {
+ state.reddit.request_fetch(item.clone());
+ }
+ }
item_page(state, uri, item).await
}
diff --git a/server/src/reddit.rs b/server/src/reddit.rs
index d203dca09245daf869b3aa942898447700ae69fb..90053ad03b1d7c8e94f325dd4ee64c2b4f7da900 100644
--- a/server/src/reddit.rs
+++ b/server/src/reddit.rs
@@ -1,4 +1,12 @@
-//! Reddit API import (async, decoupled from UI request path).
+//! Reddit API import via a single background worker (rate limits, dedup, backoff).
+
+use std::collections::{HashMap, HashSet};
+use std::sync::Arc;
+use std::time::{Duration, Instant};
+
+use reqwest::{header, Client, StatusCode};
+use serde::Deserialize;
+use tokio::sync::{mpsc, RwLock};
use crate::{
path_types::ItemId,
@@ -10,12 +18,401 @@ pub fn ensure_partial_tree(tree: &mut GlobalTree, id: &ItemId) {
tree.ensure_path(id);
}
-/// Placeholder for Reddit JSON import. Returns entity data when implemented.
-pub async fn fetch_reddit_entity(_id: &ItemId) -> Option<EntityData> {
- None
+pub struct RedditCommand {
+ pub id: ItemId,
+}
+
+#[derive(Clone)]
+pub struct RedditBroker {
+ tx: mpsc::Sender<RedditCommand>,
+}
+
+#[derive(Clone)]
+struct RedditCredentials {
+ client_id: String,
+ client_secret: String,
+}
+
+struct OAuthToken {
+ access_token: String,
+ expires_at: Instant,
+}
+
+impl RedditBroker {
+ pub fn spawn(tree: Arc<RwLock<GlobalTree>>, user_agent: &str) -> Self {
+ let (tx, rx) = mpsc::channel(100);
+
+ let mut headers = header::HeaderMap::new();
+ headers.insert(
+ header::USER_AGENT,
+ header::HeaderValue::from_str(user_agent).expect("valid user agent"),
+ );
+
+ let client = Client::builder()
+ .default_headers(headers)
+ .timeout(Duration::from_secs(15))
+ .build()
+ .expect("reqwest client");
+
+ let creds = RedditCredentials::from_env();
+ tokio::spawn(reddit_worker(rx, tree, client, creds));
+
+ Self { tx }
+ }
+
+ /// Fire-and-forget: queue a fetch; worker updates the tree when done.
+ pub fn request_fetch(&self, id: ItemId) {
+ let _ = self.tx.try_send(RedditCommand { id });
+ }
+}
+
+impl RedditCredentials {
+ fn from_env() -> Option<Self> {
+ let client_id = std::env::var("REDDIT_CLIENT_ID").ok()?;
+ let client_secret = std::env::var("REDDIT_CLIENT_SECRET").ok()?;
+ if client_id.is_empty() || client_secret.is_empty() {
+ return None;
+ }
+ Some(Self {
+ client_id,
+ client_secret,
+ })
+ }
+}
+
+pub fn default_user_agent() -> String {
+ std::env::var("REDDIT_USER_AGENT").unwrap_or_else(|_| {
+ "web:sorter2.social:v0.0.1 (by /u/sorter2)".to_string()
+ })
}
-/// Apply fetched entity data to a node (called from async worker).
-pub fn apply_entity(tree: &mut GlobalTree, id: &ItemId, data: EntityData) {
- tree.set_entity_data(id, data);
+async fn reddit_worker(
+ mut rx: mpsc::Receiver<RedditCommand>,
+ tree: Arc<RwLock<GlobalTree>>,
+ client: Client,
+ creds: Option<RedditCredentials>,
+) {
+ let mut in_flight = HashSet::new();
+ let mut recently_fetched: HashMap<ItemId, Instant> = HashMap::new();
+ let mut current_delay = Duration::from_secs(1);
+ let mut oauth: Option<OAuthToken> = None;
+ let cache_ttl = Duration::from_secs(300);
+
+ while let Some(cmd) = rx.recv().await {
+ let now = Instant::now();
+ recently_fetched.retain(|_, t| now.duration_since(*t) < cache_ttl);
+
+ if in_flight.contains(&cmd.id) || recently_fetched.contains_key(&cmd.id) {
+ continue;
+ }
+
+ in_flight.insert(cmd.id.clone());
+ let fetch_id = cmd.id.clone();
+
+ tokio::time::sleep(current_delay).await;
+
+ if let Some(c) = &creds {
+ oauth = ensure_oauth_token(&client, c, oauth.take()).await;
+ }
+
+ let token = oauth.as_ref().map(|t| t.access_token.as_str());
+ let use_oauth = token.is_some();
+
+ match do_fetch(&client, &fetch_id, use_oauth, token).await {
+ Ok(FetchOutcome::Entity(data)) => {
+ let mut w = tree.write().await;
+ w.set_entity_data(&fetch_id, data);
+ recently_fetched.insert(fetch_id.clone(), Instant::now());
+ current_delay = Duration::from_millis(600);
+ }
+ Ok(FetchOutcome::NotFound) => {
+ recently_fetched.insert(fetch_id.clone(), Instant::now());
+ }
+ Ok(FetchOutcome::RateLimited { reset_secs }) => {
+ let wait = Duration::from_secs(reset_secs.max(1));
+ tracing::warn!(
+ "Reddit rate limit for {}; sleeping {}s",
+ fetch_id,
+ wait.as_secs()
+ );
+ tokio::time::sleep(wait).await;
+ current_delay = (current_delay * 2).min(Duration::from_secs(60));
+ }
+ Err(e) => {
+ tracing::warn!("Reddit fetch failed for {}: {}", fetch_id, e);
+ current_delay = (current_delay * 2).min(Duration::from_secs(60));
+ }
+ }
+
+ in_flight.remove(&fetch_id);
+ }
+}
+
+enum FetchOutcome {
+ Entity(EntityData),
+ NotFound,
+ RateLimited { reset_secs: u64 },
+}
+
+async fn ensure_oauth_token(
+ client: &Client,
+ creds: &RedditCredentials,
+ existing: Option<OAuthToken>,
+) -> Option<OAuthToken> {
+ if let Some(t) = existing {
+ if Instant::now() < t.expires_at - Duration::from_secs(60) {
+ return Some(t);
+ }
+ }
+
+ let resp = client
+ .post("https://www.reddit.com/api/v1/access_token")
+ .basic_auth(&creds.client_id, Some(&creds.client_secret))
+ .form(&[("grant_type", "client_credentials")])
+ .send()
+ .await;
+
+ let resp = match resp {
+ Ok(r) => r,
+ Err(e) => {
+ tracing::warn!("Reddit OAuth token request failed: {e}");
+ return None;
+ }
+ };
+
+ if !resp.status().is_success() {
+ tracing::warn!("Reddit OAuth token HTTP {}", resp.status());
+ return None;
+ }
+
+ #[derive(Deserialize)]
+ struct TokenResponse {
+ access_token: String,
+ expires_in: u64,
+ }
+
+ let body: TokenResponse = match resp.json().await {
+ Ok(b) => b,
+ Err(e) => {
+ tracing::warn!("Reddit OAuth token parse failed: {e}");
+ return None;
+ }
+ };
+
+ Some(OAuthToken {
+ access_token: body.access_token,
+ expires_at: Instant::now() + Duration::from_secs(body.expires_in),
+ })
+}
+
+async fn do_fetch(
+ client: &Client,
+ id: &ItemId,
+ use_oauth: bool,
+ bearer: Option<&str>,
+) -> Result<FetchOutcome, String> {
+ let url = map_item_to_reddit_api(id, use_oauth);
+ if url.is_empty() {
+ return Ok(FetchOutcome::NotFound);
+ }
+
+ let mut req = client.get(&url);
+ if let Some(token) = bearer {
+ req = req.bearer_auth(token);
+ }
+
+ let resp = req.send().await.map_err(|e| e.to_string())?;
+
+ if resp.status() == StatusCode::TOO_MANY_REQUESTS {
+ let reset = rate_limit_reset_secs(&resp);
+ return Ok(FetchOutcome::RateLimited { reset_secs: reset });
+ }
+
+ if resp.status() == StatusCode::SERVICE_UNAVAILABLE {
+ return Err("Reddit unavailable (503)".to_string());
+ }
+
+ if !resp.status().is_success() {
+ return Ok(FetchOutcome::NotFound);
+ }
+
+ if rate_limit_remaining(&resp) == Some(0) {
+ let reset = rate_limit_reset_secs(&resp);
+ return Ok(FetchOutcome::RateLimited { reset_secs: reset });
+ }
+
+ let bytes = resp.bytes().await.map_err(|e| e.to_string())?;
+ Ok(parse_reddit_json(id, &bytes)
+ .map(FetchOutcome::Entity)
+ .unwrap_or(FetchOutcome::NotFound))
+}
+
+fn rate_limit_remaining(resp: &reqwest::Response) -> Option<u64> {
+ resp.headers()
+ .get("x-ratelimit-remaining")
+ .and_then(|v| v.to_str().ok())
+ .and_then(|s| s.parse::<f64>().ok())
+ .map(|f| f.floor() as u64)
+}
+
+fn rate_limit_reset_secs(resp: &reqwest::Response) -> u64 {
+ resp.headers()
+ .get("x-ratelimit-reset")
+ .and_then(|v| v.to_str().ok())
+ .and_then(|s| s.parse::<f64>().ok())
+ .map(|f| f.ceil() as u64)
+ .unwrap_or(5)
+}
+
+/// Map canonical item id to Reddit JSON API URL.
+pub fn map_item_to_reddit_api(id: &ItemId, oauth: bool) -> String {
+ let path = id.as_str();
+ if !path.starts_with("reddit.com/") && path != "reddit.com" {
+ return String::new();
+ }
+
+ let base = if oauth {
+ "https://oauth.reddit.com"
+ } else {
+ "https://www.reddit.com"
+ };
+
+ let segments: Vec<&str> = path.split('/').collect();
+
+ if let Some(i) = segments.iter().position(|&p| p == "comments") {
+ if segments.len() > i + 1 {
+ let api_path = segments[1..=i + 1].join("/");
+ return format!("{base}/{api_path}.json?raw_json=1");
+ }
+ }
+
+ if segments.len() == 3 && segments[1] == "r" {
+ return format!("{base}/r/{}/about.json?raw_json=1", segments[2]);
+ }
+
+ String::new()
+}
+
+fn parse_reddit_json(id: &ItemId, bytes: &[u8]) -> Option<EntityData> {
+ let v: serde_json::Value = serde_json::from_slice(bytes).ok()?;
+ let segments: Vec<&str> = id.as_str().split('/').collect();
+
+ if segments.iter().any(|&p| p == "comments") {
+ parse_post_listing(&v)
+ } else {
+ parse_subreddit_about(&v)
+ }
+}
+
+fn parse_subreddit_about(v: &serde_json::Value) -> Option<EntityData> {
+ let data = v.get("data")?;
+ let title = data
+ .get("title")
+ .or_else(|| data.get("display_name"))
+ .and_then(|t| t.as_str())?
+ .to_string();
+ let body_html = data
+ .get("public_description_html")
+ .or_else(|| data.get("public_description"))
+ .and_then(|t| t.as_str())
+ .map(|s| s.to_string());
+ let thumb_url = data
+ .get("icon_img")
+ .or_else(|| data.get("community_icon"))
+ .and_then(|t| t.as_str())
+ .filter(|s| !s.is_empty())
+ .map(|s| s.to_string());
+
+ Some(EntityData {
+ title,
+ author: None,
+ body_html,
+ thumb_url,
+ })
+}
+
+fn parse_post_listing(v:
… preview truncated; 4,570 characters omittedHardlinks — judgments / attempts / prompt
judgments
attempts
Prompt text is loaded only by the download route.