You are a constitutional council ranking individual git commits for ownership allocation. Compare these two commits. Decide which contributed more lasting value to the project. Judge substance, not spectacle: - Prefer correct, lasting design and real bugfixes over churn, formatting, renames, or generated noise. - Prefer clarity and necessity over sheer line count. A small precise change can beat a large diffuse one. - Do not favor a side merely because its patch is longer or noisier. - Weight what the change does for the project, not the contributor's name. Return ONLY a JSON object: {"winner": "A" or "B", "ratio": "N:M", "explanation": "..."} The explanation must cite concrete differences in the patches (1-3 sentences). Side A — contributor: tommy-mor Side A — commit message: [8ec3e9c4] nice Side A — unified diff (full patch): diff --git a/deploy/open-webui/README.md b/deploy/open-webui/README.md new file mode 100644 index 0000000000000000000000000000000000000000..0f1317f745f31530fa687c8a35e5e3a726d1dda2 --- /dev/null +++ b/deploy/open-webui/README.md @@ -0,0 +1,71 @@ +# Open WebUI on Fly.io (minimal, OpenRouter) + +You do **not** need to host your own model. [OpenRouter](https://openrouter.ai/) runs the LLMs; this Fly app only runs the Open WebUI shell. + +### Disk vs RAM + +The container image still includes ML libraries on **disk** (upstream wheels). These settings target **runtime memory**: chat and API proxy use remote inference, and `RAG_EMBEDDING_ENGINE=openai` plus `AUDIO_STT_ENGINE=openai` tell Open WebUI to call your configured OpenAI-compatible API for embeddings and speech-to-text instead of loading local SentenceTransformers / Whisper models into RAM (see Open WebUI [performance](https://docs.openwebui.com/troubleshooting/performance/) notes on embedding offload). + +You may still see higher RSS if you use features that pull in other local code paths. For a fresh volume, the `[env]` values apply on first boot; if you already ran Open WebUI, stored admin settings can override some env vars ([PersistentConfig](https://docs.openwebui.com/reference/env-configuration/) — adjust in Admin or reset the volume for a clean slate). + +## Prerequisites + +- [Fly CLI](https://fly.io/docs/hands-on/install-flyctl/) and `fly auth login` +- An OpenRouter API key (`sk-or-…` from [openrouter.ai/keys](https://openrouter.ai/keys)) + +## One-time setup + +1. Edit `fly.toml` and set `app = "your-unique-name"` (globally unique on Fly). + +2. Create the app (if it does not exist): + + ```bash + fly apps create your-unique-name + ``` + +3. Create a volume for SQLite and uploads (region must match `primary_region` in `fly.toml`): + + ```bash + fly volumes create open_webui_data --region iad --size 3 + ``` + +4. Set secrets: + + ```bash + fly secrets set OPENAI_API_KEY="sk-or-..." \ + WEBUI_SECRET_KEY="$(openssl rand -hex 32)" \ + WEBUI_URL="https://your-unique-name.fly.dev" + ``` + + Optional but convenient: create the first admin in one step (disables open signup on first boot): + + ```bash + fly secrets set WEBUI_ADMIN_EMAIL="you@example.com" WEBUI_ADMIN_PASSWORD='strong-password-here' + ``` + +5. Deploy: + + ```bash + cd deploy/open-webui && fly deploy + ``` + +Open `https://your-unique-name.fly.dev`. In the UI, pick a model served by OpenRouter (IDs like `openai/gpt-4o`, `anthropic/claude-3.5-sonnet`, etc.). + +## If the machine runs out of memory + +The `fly.toml` env is tuned to avoid local embedding/STT models in RAM; if it still OOMs (large chats, document uploads, many users), scale up: + +```bash +fly scale memory 2048 +# or 4096 if needed +``` + +## When you *would* self-host a model + +Only if you want **local / private inference** (no third-party API). That usually means **Ollama** or similar on a GPU-capable host, not this minimal Fly setup. For “just chat via API,” OpenRouter (or direct OpenAI) is enough. + +## Notes + +- `OPENAI_API_BASE_URL` points at OpenRouter’s OpenAI-compatible API; `OPENAI_API_KEY` is your OpenRouter key. +- After the first run, some settings are stored in the volume under `/app/backend/data` (see Open WebUI docs on `PersistentConfig`). +- Pin the image tag instead of `:main` in `fly.toml` if you want reproducible deploys. diff --git a/deploy/open-webui/fly.toml b/deploy/open-webui/fly.toml new file mode 100644 index 0000000000000000000000000000000000000000..684ae6d108e6a90906ef95c5075b468fe8280134 --- /dev/null +++ b/deploy/open-webui/fly.toml @@ -0,0 +1,59 @@ +# Open WebUI on Fly.io — OpenRouter only (no Ollama). +# Runtime RAM: RAG_EMBEDDING_ENGINE + AUDIO_STT_ENGINE use your OpenRouter HTTP API so local +# SentenceTransformers / Whisper weights are not loaded into memory (torch may still exist on disk in the image). +# Replace `app` with your Fly app name, then: +# fly volumes create open_webui_data --region --size 3 +# fly secrets set OPENAI_API_KEY=sk-or-... WEBUI_SECRET_KEY=$(openssl rand -hex 32) +# fly secrets set WEBUI_URL=https://.fly.dev +# Optional (recommended): headless admin, signup disabled automatically: +# fly secrets set WEBUI_ADMIN_EMAIL=you@example.com WEBUI_ADMIN_PASSWORD='...' +# fly deploy + +app = "open-webui" +primary_region = "iad" + +[build] + image = "ghcr.io/open-webui/open-webui:main" + +[env] + OPENAI_API_BASE_URL = "https://openrouter.ai/api/v1" + ENABLE_OLLAMA_API = "false" + # One worker so we do not duplicate in-RAM state (see Open WebUI scaling docs). + UVICORN_WORKERS = "1" + # Remote embeddings via OpenAI-compatible API (OpenRouter) — avoids ~500MB+ local embedding model in RAM. + RAG_EMBEDDING_ENGINE = "openai" + # Remote speech-to-text when using voice — avoids loading faster-whisper into RAM. + AUDIO_STT_ENGINE = "openai" + # Speeds up model list with OpenRouter’s large catalog (small in-memory cache). + ENABLE_BASE_MODELS_CACHE = "true" + +[[mounts]] + source = "open_webui_data" + destination = "/app/backend/data" + +[[services]] + internal_port = 8080 + protocol = "tcp" + + [[services.ports]] + port = 80 + handlers = ["http"] + force_https = true + + [[services.ports]] + port = 443 + handlers = ["tls", "http"] + + [services.concurrency] + type = "connections" + hard_limit = 250 + soft_limit = 200 + + [[services.http_checks]] + interval = "15s" + timeout = "5s" + grace_period = "60s" + method = "GET" + path = "/health" + protocol = "http" + tls_skip_verify = false Side B — contributor: tommy-mor Side B — commit message: [8f69c309] Require Reddit OAuth when credentials are set and refresh on 401/403. Avoid falling back to the public www.reddit.com API from cloud IPs, which returns Reddit's network-security block page. Also pin SORTER2_BASE_URL in fly.toml. Co-authored-by: Cursor Side B — unified diff (full patch): diff --git a/fly.toml b/fly.toml index f0c6a39f643c204987debc177234d95a8ef44b65..ca7e0088a7efd7d58d29d808f8d89080b2bff233 100644 --- a/fly.toml +++ b/fly.toml @@ -5,6 +5,7 @@ primary_region = "iad" dockerfile = "Dockerfile" [env] + SORTER2_BASE_URL = "https://reddit.sorter.social" SORTER2_DATA_DIR = "/data" SORTER2_EVENT_LOG = "/data/events.jsonl" PORT = "8080" diff --git a/server/src/reddit.rs b/server/src/reddit.rs index a874814f8927192ee62cab2d0db1efd27dcd57b7..f409764c1e1f36216f1b08107043c2eab905694c 100644 --- a/server/src/reddit.rs +++ b/server/src/reddit.rs @@ -283,26 +283,35 @@ async fn reddit_worker( ); tokio::time::sleep(current_delay).await; - if let Some(c) = &creds { - oauth = ensure_oauth_token(&client, &oauth_token_base, c, oauth.take()).await; - } - - let token = oauth.as_ref().map(|t| t.access_token.as_str()); - let fetch_base = if token.is_some() { - tracing::debug!( - item = %fetch_id, - base = %oauth_api_base, - "reddit fetch using OAuth bearer" - ); - &oauth_api_base - } else { - &api_base - }; - let url = match kind { - FetchKind::SelfEntity => map_item_to_reddit_api(&fetch_id, fetch_base), - FetchKind::Children => map_children_url(&fetch_id, fetch_base), + let outcome = match &creds { + Some(c) => { + // OAuth is required when credentials are configured — never fall + // back to the public www.reddit.com JSON endpoints (cloud IPs + // get blocked with a 403 HTML interstitial). + fetch_with_oauth( + &client, + &oauth_token_base, + &oauth_api_base, + c, + &mut oauth, + &fetch_id, + kind, + ) + .await + } + None => { + let url = match kind { + FetchKind::SelfEntity => map_item_to_reddit_api(&fetch_id, &api_base), + FetchKind::Children => map_children_url(&fetch_id, &api_base), + }; + match do_fetch(&client, &url, &fetch_id, None).await { + Ok(FetchOutcome::AuthRejected { status, detail }) => { + Err(format!("Reddit API {status}: {detail}")) + } + other => other, + } + } }; - let outcome = do_fetch(&client, &url, &fetch_id, token).await; match outcome { Ok(FetchOutcome::Payload(payload)) => { @@ -342,6 +351,12 @@ async fn reddit_worker( current_delay = (current_delay * 2).min(Duration::from_secs(60)); notify(done, FetchJobResult::RateLimited { reset_secs }); } + Ok(FetchOutcome::AuthRejected { status, detail }) => { + let e = format!("Reddit API {status}: {detail}"); + tracing::warn!(item = %fetch_id, err = %e, "reddit fetch auth rejected"); + current_delay = (current_delay * 2).min(Duration::from_secs(60)); + notify(done, FetchJobResult::Failed(e)); + } Err(e) => { tracing::warn!(item = %fetch_id, err = %e, "reddit fetch failed"); current_delay = (current_delay * 2).min(Duration::from_secs(60)); @@ -357,6 +372,60 @@ enum FetchOutcome { Payload(Value), NotFound, RateLimited { reset_secs: u64 }, + /// Bearer rejected — caller should drop the cached token and retry once. + AuthRejected { status: StatusCode, detail: String }, +} + +async fn fetch_with_oauth( + client: &Client, + oauth_token_base: &str, + oauth_api_base: &str, + creds: &RedditCredentials, + oauth: &mut Option, + fetch_id: &ItemId, + kind: FetchKind, +) -> Result { + for attempt in 0..2 { + let force_refresh = attempt > 0; + *oauth = Some( + ensure_oauth_token(client, oauth_token_base, creds, oauth.take(), force_refresh) + .await?, + ); + let token = oauth + .as_ref() + .expect("token set above") + .access_token + .clone(); + + tracing::debug!( + item = %fetch_id, + base = %oauth_api_base, + attempt, + "reddit fetch using OAuth bearer" + ); + + let url = match kind { + FetchKind::SelfEntity => map_item_to_reddit_api(fetch_id, oauth_api_base), + FetchKind::Children => map_children_url(fetch_id, oauth_api_base), + }; + match do_fetch(client, &url, fetch_id, Some(&token)).await? { + FetchOutcome::AuthRejected { status, detail } if attempt == 0 => { + tracing::warn!( + item = %fetch_id, + %status, + %detail, + "reddit OAuth rejected; refreshing token and retrying" + ); + *oauth = None; + continue; + } + FetchOutcome::AuthRejected { status, detail } => { + return Err(format!("Reddit API {status}: {detail}")); + } + other => return Ok(other), + } + } + unreachable!("loop always returns") } async fn ensure_oauth_token( @@ -364,35 +433,35 @@ async fn ensure_oauth_token( oauth_base: &str, creds: &RedditCredentials, existing: Option, -) -> Option { - if let Some(t) = existing { - if Instant::now() < t.expires_at - Duration::from_secs(60) { - tracing::debug!("reddit OAuth token still valid"); - return Some(t); + force_refresh: bool, +) -> Result { + if !force_refresh { + if let Some(t) = existing { + if Instant::now() < t.expires_at - Duration::from_secs(60) { + tracing::debug!("reddit OAuth token still valid"); + return Ok(t); + } } } let url = format!("{}/api/v1/access_token", oauth_base.trim_end_matches('/')); - tracing::debug!(%url, "reddit OAuth token request"); + tracing::debug!(%url, force_refresh, "reddit OAuth token request"); let resp = client .post(&url) .basic_auth(&creds.client_id, Some(&creds.client_secret)) .form(&[("grant_type", "client_credentials")]) .send() - .await; - - let resp = match resp { - Ok(r) => r, - Err(e) => { - tracing::warn!("reddit OAuth token request failed: {e}"); - return None; - } - }; + .await + .map_err(|e| format!("Reddit OAuth token request failed: {e}"))?; if !resp.status().is_success() { - tracing::warn!("reddit OAuth token HTTP {}", resp.status()); - return None; + let status = resp.status(); + let body = resp.text().await.unwrap_or_default(); + return Err(format!( + "Reddit OAuth token HTTP {status}: {}", + truncate_for_error(&body) + )); } #[derive(Deserialize)] @@ -401,21 +470,40 @@ async fn ensure_oauth_token( expires_in: u64, } - let body: TokenResponse = match resp.json().await { - Ok(b) => b, - Err(e) => { - tracing::warn!("reddit OAuth token parse failed: {e}"); - return None; - } - }; + let body: TokenResponse = resp + .json() + .await + .map_err(|e| format!("Reddit OAuth token parse failed: {e}"))?; - tracing::debug!(expires_in = body.expires_in, "reddit OAuth token acquired"); - Some(OAuthToken { + tracing::info!(expires_in = body.expires_in, "reddit OAuth token acquired"); + Ok(OAuthToken { access_token: body.access_token, expires_at: Instant::now() + Duration::from_secs(body.expires_in), }) } +fn truncate_for_error(body: &str) -> String { + let compact: String = body.split_whitespace().collect::>().join(" "); + if compact.is_empty() { + return "(empty body)".into(); + } + // Prefer the human-readable block message over dumping Reddit's CSS. + if let Some(idx) = compact.find("You've been blocked") { + let slice: String = compact.chars().skip(idx).take(160).collect(); + return if compact.chars().count() > idx + 160 { + format!("{slice}…") + } else { + slice + }; + } + let chars: String = compact.chars().take(200).collect(); + if compact.chars().count() > 200 { + format!("{chars}…") + } else { + chars + } +} + async fn do_fetch( client: &Client, url: &str, @@ -460,15 +548,16 @@ async fn do_fetch( if !status.is_success() { let body = resp.text().await.unwrap_or_default(); + let detail = truncate_for_error(&body); tracing::debug!( item = %id, %status, body_len = body.len(), - body_prefix = %body.chars().take(240).collect::(), + %detail, "reddit non-success body" ); if status == StatusCode::FORBIDDEN || status == StatusCode::UNAUTHORIZED { - return Err(format!("Reddit API {status}: {body}")); + return Ok(FetchOutcome::AuthRejected { status, detail }); } return Ok(FetchOutcome::NotFound); } @@ -716,6 +805,15 @@ fn reddit_direct_image_url(url: &str) -> bool { mod tests { use super::*; + #[test] + fn truncate_error_prefers_block_message() { + let html = r#"
You've been blocked by network security. To continue, log in
"#; + let msg = truncate_for_error(html); + assert!(msg.starts_with("You've been blocked")); + assert!(msg.len() < 200); + assert!(!msg.contains(".x{color")); + } + #[test] fn map_subreddit_about_url() { let id = ItemId::from_url("https://reddit.com/r/rust").unwrap();