comparison · c_16438843de8f (tommy-mor) vs c_7a129e904906 (tommy-mor)
Side A adds new files (Dockerfile, deps.edn, event_log.rs, views.rs, fly.toml) that are not wired into any build system or referenced by existing code, appearing as unintegrated scaffolding/dead code dropped at repo root. Side B is a substantive, integrated feature: it adds a real-time SSE audit dashboard, fixes an object-hash dedup bug, gates contested rankings on OPENROUTER_API_KEY with tests, adds a CI/CD deploy workflow, and includes corresponding integration and unit tests validating the new behavior.
B ships production deployment (Fly Dockerfile/fly.toml + main-branch CI), a full live /watch audit UI with SSE status/progress, multi-repo/contributor production roots, epoch-loop reliability fixes, and matching tests. A only drops early Rust seed scaffolding (event_log/views + Dockerfile/deps) without wiring, tests, or evidence it became the lasting path.
Side B delivers substantial, lasting functionality by adding a production deployment pipeline (GitHub Actions, Docker, Fly.io), a live audit/status system with structured SSE events, a new /watch dashboard, process-state tracking, authenticated Git access, and accompanying integration/unit tests. Side A mainly introduces infrastructure and persistence helpers (Dockerfile, event log, view counter, deployment config), which are useful but narrower in scope and less integrated into the application's core behavior.
comparison · c_7a129e904906 (tommy-mor) vs c_6864b1ca8ce6 (tommy-mor)
Side A adds substantive infrastructure and functionality: CI/CD deployment pipeline, Dockerfile, fly.toml, GitHub token auth for private git mirroring, a real /watch dashboard with SSE-driven audit trail and process state, plus new tests covering ranking edge cases (single contributor, missing API key) and audit event broadcasting. Side B is purely CSS polish (padding, border-radius, focus outlines, color tweaks) across two theme files with no new tests, functionality, or lasting architectural value.
A ships production deployment (Dockerfile, fly.toml, CI deploy), a real /watch audit UI with SSE progress/state APIs, multi-repo/contributor config, git auth and emission reliability fixes, plus tests—lasting product and ops capability. B only restyles existing vote/pin controls in two CSS themes (padding, borders, focus rings) with no behavior, API, or architecture change.
Side A adds substantial new functionality and operational infrastructure: production deployment (Docker, Fly, GitHub Actions), a live audit dashboard with SSE event streaming, status APIs, authenticated Git access, improved repository discovery efficiency, startup/retry logic, and accompanying integration/unit tests. Side B is almost entirely CSS refinements for the voting UI—improving layout, styling, focus states, and accessibility—but it does not introduce comparable core behavior or architecture.
comparison · c_e4fb43f04791 (tommy-mor) vs c_7a129e904906 (tommy-mor)
Side B ships substantial, functional infrastructure and features: a deployment pipeline (Dockerfile, fly.toml, GitHub Actions), a real /watch dashboard with live SSE audit events, security-relevant fixes (git auth headers, decimal context, single-contributor bypass, OpenRouter key guard), and new tests covering these behaviors. Side A is a small CSS/markup cleanup (removing a wrapper div and restyling rank numbers) that is tidy but of minor, cosmetic scope compared to B's operational and functional impact.
B ships production deployment (Docker, fly.toml, CI test-and-deploy), expands multi-repo/contributor configuration, and adds a real audited live /watch dashboard with SSE progress, status API, and tests—core lasting product and ops value. A only unwraps a vote-compare shell div and tweaks ontology ranking-list CSS counter styling, which is small presentational churn.
Side B adds substantial, lasting infrastructure and application functionality: production deployment (Dockerfile, Fly.io, GitHub Actions), a live audit dashboard with SSE-backed status updates, a new status API, improved Git authentication/configuration, deduplicated Git object verification, runtime error handling, and accompanying integration/unit tests. Side A mainly removes a wrapper element from the HTML template and adjusts CSS for ranking list styling, with only cosmetic/layout impact and no comparable functional improvement.
comparison · c_55f1cdf12e22 (tommy-mor) vs c_7a129e904906 (tommy-mor)
Side A implements a coherent new feature (invite links) end-to-end: server events/reducer state, RPC handlers, CLI commands, HTML routes, and a dedicated integration test file, all wired through existing patterns with real design tradeoffs (TTL, exhaustion, redemption tying into OAuth flow). Side B is largely deployment/infra glue (Dockerfile, fly.toml, CI) plus a UI dashboard and audit-broadcast plumbing that, while functional, is more operational scaffolding and cosmetic CSS than durable core-domain logic, and it hardcodes org-specific repo/contributor config that is brittle and less generally reusable.
A lands a full invite lifecycle (mint RPC, /join redemption into grants, multi-cap RoomGrant, RoomAudit, CLI, reducer/timeline types) with a dedicated invites integration test—durable product capability. B’s deploy pipeline, /watch SSE audit UI, and production repo/contributor config are operationally important but mostly wire the existing constitution loop for production rather than adding comparable core domain behavior.
Side A implements substantial new project functionality: an invite-based room access flow spanning server, CLI, RPC, authentication, state management, API types, routing, and end-to-end integration tests, while also extending grants to multiple capabilities and adding room audit support. Side B mainly adds deployment infrastructure, a production dashboard, SSE audit/status reporting, and CI/CD configuration; valuable operationally, but it contributes less core application behavior than the end-user invite and permission features in Side A.
comparison · c_11ce057e37af (tommy-mor) vs c_7a129e904906 (tommy-mor)
Side B ships a coherent, tested feature set (deploy pipeline, Dockerfile, fly.toml, GitHub auth for git mirroring, a live /watch dashboard with SSE audit events, and single-contributor fast-path fix for rank_commits) backed by unit tests and integration test updates, giving durable infrastructure and correctness value. Side A is a solid, well-tested parser improvement (deterministic tokens, prose URL linkification, braced-body enforcement) but is narrower in scope and mostly internal refactoring/feature polish within one module.
B delivers lasting operational value: production deploy (Docker/Fly/CI), multi-repo contributor config, emission retry, and a real audited /watch SSE progress surface with status API and tests—foundational for running the constitution itself. A is strong design (deterministic typed BlockMasker, prose ref tokenizer, braced body rules, linkify), but it is product/DSL polish plus some formatting churn, not the same system-level necessity.
Side A makes substantive parser and rendering improvements: it introduces typed deterministic block masking, a prose item-reference tokenizer that correctly handles raw URLs, punctuation, newlines, and code fences, enforces braced DSL item bodies, updates linkification accordingly, and adds focused tests for these behaviors. Side B adds valuable deployment infrastructure and a live audit dashboard (Docker, CI/CD, SSE status/UI, production config), but much of its impact is operational rather than improving the project's core parsing and data model, so its lasting contribution is somewhat smaller.
comparison · c_9bced108c8aa (tommy-mor) vs c_7a129e904906 (tommy-mor)
B ships production deployment infrastructure (Dockerfile, fly.toml, CI workflow) plus real correctness fixes (git credential injection for private mirrors, object-hash dedup by location instead of by ref, requiring OPENROUTER_API_KEY only when contested, safer SSE client removal) and a functional live-audit dashboard with tests. A is a clean, well-tested URL canonicalization graph, but it's a narrower, self-contained feature versus B's broader operational and correctness improvements that let the whole system actually run and be observed in production.
A introduces a full semantic URL DFA (graph engine, builder with link validation, parsing/normalization, and dense behavioral tests) that is core lasting product design. B mainly ships production wiring (Docker/Fly/CI), multi-repo config, and an audit/watch SSE UI—valuable operationally but more scaffolding and presentation around an existing emission pipeline than new foundational logic.
Side A adds a substantial new URL canonicalization subsystem: a graph-based traversal engine with declarative graph builder, URL parsing/normalization, generic fallback logic, breadcrumbs, and extensive tests covering Reddit, YouTube, encoding, query stripping, and traversal behavior. Side B mainly adds deployment infrastructure, production configuration, a live monitoring UI/SSE audit stream, and some robustness improvements (e.g. GitHub token support and deduplicated object verification), which are valuable operational enhancements but less foundational to the project's core behavior than the new canonicalization engine.
comparison · c_df12ba3b70a8 (tommy-mor) vs c_7a129e904906 (tommy-mor)
Side A fixes a real bug (empty external garden index due to bogus parent), cleanly restructures the GitHub resolver into a resolvers/ module with a proper card schema/renderer, and adds targeted unit/integration tests validating the new behavior. Side B is largely infra/deploy plumbing (Dockerfile, fly.toml, CI) plus a sizable dashboard/SSE feature for a separate constitution.py service, which is useful operationally but is less about core product correctness and mixes config, UI, and scattered test tweaks.
B ships the constitution as a running, auditable production system: Fly/Docker/CI deploy path, expanded real repo/contributor roots, resilient epoch execution, /api/status, and a tested /watch SSE progress UI—foundational lasting infrastructure. A’s /- host-index bugfix, resolvers/ card schema, and GitHub rich rendering are strong product work, but more incremental on slug.social than standing up the live ownership process itself.
Side A fixes a real functional bug by replacing the bogus `https://.` parent lookup for the external garden index with `external_root_host_items`, ensuring external roots are discovered even from implicit child edges, and adds tests for that behavior. It also introduces a reusable resolver architecture (`server/src/resolvers/`), rich GitHub import card rendering via `render_item_body_in_scope`, and integration/unit tests, whereas Side B is primarily deployment and observability infrastructure (Docker, Fly, GitHub Actions, SSE dashboard, status APIs) that improves operations but less directly changes the project's core behavior.
comparison · c_77729db919ab (tommy-mor) vs c_7a129e904906 (tommy-mor)
Side A is a substantive, well-tested refactor of URL canonicalization into a composable rule engine (url_rules module) that fixes real correctness issues (scheme-qualified ItemId, reddit/youtube normalization, parent/breadcrumb logic) with extensive new unit tests. Side B is largely infra/ops scaffolding (Dockerfile, fly.toml, CI) plus a dashboard UI and audit-broadcast plumbing that adds operational value but is less architecturally durable and more application-specific/generated-feeling (large CSS/JS blob, repo-list config tweaks).
A delivers lasting core design: a composable url_rules engine, scheme-full canonical ItemIds, and correct parent/breadcrumb hierarchy (e.g. skipping phantom /comments nodes) wired through identity, Reddit mapping, and projection apply. B’s deploy path, status API, epoch/OpenRouter guards, and multi-repo config are real operational value, but a large share is dashboard CSS/JS surface and infra wiring rather than durable domain structure.
Side A introduces a substantive URL canonicalization architecture by extracting normalization into composable `url_rules` modules, making `ItemId` consistently use canonical HTTPS URLs, fixing parent/breadcrumb handling, and updating projection/event parsing to canonicalize stored IDs. Side B adds valuable deployment infrastructure, a live audit dashboard, SSE progress reporting, and CI/CD, but those are primarily operational features rather than core data-model improvements; the canonical identity changes in A have broader long-term impact on correctness and consistency.
comparison · c_4772ee88dbe3 (tommy-mor) vs c_7a129e904906 (tommy-mor)
Side A is a reasonable internal refactor (async settlement worker, cache separation, removal of demo scaffolding) but touches a relatively contained area with modest complexity risk (e.g., worker batching without careful backpressure/error semantics). Side B ships substantial production infrastructure (Dockerfile, fly.toml, CI/CD deploy workflow), fixes real correctness issues (duplicate-ref commit hashing, single-contributor ranking shortcut, missing OPENROUTER_API_KEY guard), and adds a genuinely useful live-audit /watch UI with corresponding tests, representing broader lasting value to the deployed system.
B lands production authority (Fly Dockerfile/fly.toml, main-branch test-and-deploy CI), multi-repo/contributor roots, and a real auditable /watch+SSE/status pipeline with epoch-loop retry and ranking guards—core operational value. A’s lasting piece is the settlement worker plus cached read-path rankings and demo-counter removal, but that is a narrower sorter2 refactor versus making the constitution process deployable and observable.
Side A makes substantive architectural changes to the server: it removes the demo-only counter and related event type, introduces a settlement worker that batches vote persistence and ranking recomputation, adds cached ranking reads (`ranked_items_cached`) to avoid unnecessary recomputation, and switches UI paths from write locks to read locks for rendering. Side B adds valuable deployment infrastructure, a live audit dashboard, SSE status reporting, Docker/Fly configuration, and CI, but much of its patch is operational/UI surface rather than improving the project's core behavior, so A provides the stronger lasting technical foundation.
comparison · c_7a129e904906 (tommy-mor) vs c_97611919bf0b (tommy-mor)
B is a substantive, correctness-focused type refactor (CanonicalItemUrl -> ItemId) that touches the core reducer/ranking/RPC layers with real design tradeoffs (enum variants, parent/tilde logic, serde) and includes updated tests, representing durable architectural investment. A is mostly deploy/ops scaffolding (Dockerfile, fly.toml, CI) plus a UI/audit-feed feature layered on existing code—useful but more additive/operational than foundational, with much of the diff being CSS/JS glue rather than core logic changes.
A ships production deployment (Fly/Docker/CI), a real /watch audit UI with SSE progress, epoch failure recovery, and multi-repo/contributor config—turning the constitution into a live, operable system. B is mostly a CanonicalItemUrl→ItemId migration across the slug server; the enum still largely wraps strings plus Opaque fallbacks, so it is valuable typing cleanup but less net product durability than A.
Side B introduces a durable architectural change by replacing the string-based `CanonicalItemUrl` with a structured `ItemId` type, propagating it through the reducer, ranking, routing, RPC, and tests while extracting shared normalization into `item_wire.rs`. Side A adds valuable deployment infrastructure, production configuration, and a live audit dashboard with SSE/status APIs, but much of its impact is operational and UI-oriented, whereas B changes the project's core identity model in a way that simplifies future development and reduces reliance on ad hoc string handling.
comparison · c_7a129e904906 (tommy-mor) vs c_3f420a1f5aa1 (tommy-mor)
Side A ships production deployment infrastructure (Dockerfile, fly.toml, CI test+deploy workflow), fixes real correctness issues (git auth via GITHUB_TOKEN for private mirrors, requiring OPENROUTER_API_KEY before contested ranking, deduping redundant object-hash checks), and adds a functioning audited dashboard backed by new unit/integration tests. Side B is a solid but narrower fix (private-room URL prefixing across many RPC call sites plus a cookie-based theme switch) with good targeted tests, but it lacks the operational/deployment impact and breadth of correctness fixes found in A.
A ships production deployment (Dockerfile, fly.toml, main-branch CI), multi-repo roots/contributors, GitHub-auth mirrors, epoch retry, and a tested live /watch audit SSE surface that makes the constitutional process operable and observable. B improves theme persistence via cookies and correctly room-prefixes many RPC/web URLs, which is real lasting API fix work, but it is incremental product polish relative to A’s end-to-end deployable constitution runtime.
Side A delivers substantial operational and product functionality: it adds production deployment (Dockerfile, Fly config, GitHub Actions), a live audit dashboard with SSE-backed progress and status APIs, improves emission processing and retry behavior, optimizes Git object verification to avoid redundant work, and adds tests for the new behavior. Side B is a useful but narrower feature set, primarily introducing persistent theme handling across pages and correct room-aware URL generation for RPC/web responses, with corresponding plumbing and helper functions.
comparison · c_7a129e904906 (tommy-mor) vs c_afa638171cf7 (tommy-mor)
Side B implements a real, self-contained feature (multi-provider OAuth linking with UUID-as-canonical-identity, conflict handling, privacy of linked providers) backed by new server logic, storage schema changes, and updated tests/mocks—clear lasting design value. Side A is mostly deployment/infra plumbing (Dockerfile, fly.toml, CI) plus a large dashboard UI/CSS/JS addition that, while functional, is more scaffolding and observability sugar than core product logic, with less durable architectural weight.
A ships production authority (Dockerfile, fly.toml, main-branch test+deploy) and lasting constitution behavior: live audit SSE/state, /watch UI, /api/status, emission retry, multi-repo/contributor roots, and focused tests—core to an auditable ownership system. B is solid identity design (UUID-canonical multi-provider linking, Reddit OAuth, private linked-providers, trust-weight batch fix) but is a scoped auth feature versus A’s end-to-end deployable process surface.
Side B makes a durable architectural change by making UUIDs the canonical identity, refactoring OAuth into a provider-agnostic linking model, adding Reddit OAuth support, handling account-link conflicts, updating routing, projection logic, storage queries, and test infrastructure. Side A adds valuable deployment, monitoring, and operational improvements (Fly deployment, live SSE audit dashboard, status API, and production workflow), but much of its impact is operational/UI-focused rather than changing the project's core identity and authentication model.
comparison · c_2dc96aace098 (tommy-mor) vs c_7a129e904906 (tommy-mor)
Side A vendors an entire generic 'durable' RocksDB collections crate (with unused Entry/nested-collection API, README, docs, LICENSE, benchmarks, examples) just to back a simple string->string JSON store, adding large dependency/boilerplate weight disproportionate to the actual server-side change (entity_store.rs + streaming replay), much of which reads as generated filler. Side B's patch is tightly scoped to real, used functionality: CI/CD deploy pipeline, Dockerfile/fly.toml for production, a working live SSE audit dashboard, and concrete correctness fixes (require OPENROUTER_API_KEY for contested rankings, dedupe redundant git-object hashing, GITHUB_TOKEN auth), all backed by updated tests exercising the new behavior.
A’s lasting value is architectural: entity JSON moves out of in-memory GlobalTree into a RocksDB-backed EntityStore, and event_log gains true line-at-a-time replay so startup no longer materializes the full log/payloads in RAM—concrete reducer/state/reddit call-site changes that change scalability. B is real product/ops work (Fly/Dockerfile/CI deploy, /watch audit SSE, production repo/contributor roots, emission retry), but it is mostly deployment, UI, and configuration on an existing process rather than a deeper data-plane redesign.
Side A makes substantive architectural changes: it introduces a RocksDB-backed durable storage crate, moves large Reddit JSON payloads out of the in-memory tree into a persistent entity store, and replaces whole-log loading with streaming event replay to reduce startup memory, updating replay and integration tests accordingly. Side B primarily adds deployment infrastructure, a live dashboard, SSE status/audit UI, and CI/CD configuration, which improves operations and visibility but has much less impact on the project's core functionality and architecture.
comparison · c_7a129e904906 (tommy-mor) vs c_af08bd851e49 (tommy-mor)
Commit A ships a complete, coherent production deployment path (Dockerfile, fly.toml, CI/CD workflow, GitHub App auth for git mirroring, expanded repo/contributor config, a real live audit/status dashboard with SSE) plus corresponding tests, closing a previously flagged TODO. Commit B, despite its 'refactor' label, is a substantial feature addition (pairwise vote comparison page, bridge-pair selection algorithm, ItemId normalization) with decent test coverage, but is narrower in scope and less clearly load-bearing for overall project operability than A's deployment infrastructure and observability work.
A ships production authority (Fly/Docker/CI), multi-repo contributor roots, epoch retry safety, and a tested live /watch audit SSE surface that makes the constitution operable and auditable. B’s bridge-first pair selection, ItemId::from_storage normalization, and in-place vote morphs are solid core design, but they refine an existing sorter vote path rather than standing up the deployed economic process.
Side A delivers a substantial production capability: it adds deployment infrastructure (Dockerfile, Fly.io config, GitHub Actions), a live audit dashboard with SSE-backed progress/status APIs, startup/retry improvements, authenticated GitHub access, and fixes repository discovery by deduplicating commit object verification, all backed by new integration and unit tests. Side B contains useful UI and architecture work for pairwise voting (new vote page, pair-selection logic, ID normalization, and incremental morph updates), but its impact is narrower and primarily focused on one feature area rather than end-to-end operational capability.
comparison · c_7a129e904906 (tommy-mor) vs c_2595b6007624 (tommy-mor)
Side A ships production-critical infrastructure (Dockerfile, CI/CD with test gating, fly.toml) plus a working live audit dashboard, and includes concrete bugfixes (OPENROUTER_API_KEY guard for contested rankings, GITHUB_TOKEN auth for private git fetches, dedup of object-hash checks, epoch retry logic) all backed by new passing tests. Side B is a large, self-described 'first pass' rewrite of the entire server API into an RPC batch model, which is architecturally ambitious but admittedly unfinished, deletes existing test coverage (e.g. debug_query_params.rs) without full replacement, and carries higher risk of incompleteness/bugs typical of a first-draft refactor.
B redesigns the core product surface: it replaces many REST handlers with a batch RPC API, splits rooms (permission boundary) from forum thread tags in events/reducer state, and threads that model through CLI, types, and tests. A is high-value operational work (Fly deploy, CI gate, /watch audit UI, emission retries and status SSE), but it overlays constitution.py rather than changing enduring domain architecture, so B edges it on lasting design despite A’s polish.
Side B delivers a substantial architectural change by consolidating many REST endpoints into a typed RPC batch API, refactoring shared validation into a dedicated module, introducing room-scoped operations, and updating the CLI, server, event model, and tests to match. Side A adds valuable deployment infrastructure, a live audit dashboard, SSE progress reporting, and operational improvements, but much of its impact is in observability and deployment rather than the project's core API and data model.
comparison · c_7a129e904906 (tommy-mor) vs c_94135a1c4c58 (tommy-mor)
Side B fixes concrete security bugs (votes silently falling back to an anonymous actor instead of failing closed, missing Secure cookie flag, mock-OAuth escape hatch reachable in production, open-redirect edge case, and pinning the durable dependency by rev instead of a mutable branch) with focused tests for each fix. Side A is a large, mostly additive commit (Dockerfile, fly.toml, CI/deploy workflow, a new dashboard UI/CSS/JS) that does contain a couple of real fixes (per-repo object-hash dedup, fail-closed OPENROUTER_API_KEY check), but the bulk of the diff is spectacle-heavy UI/infra rather than correctness-critical logic.
A ships production authority (Docker/Fly/CI deploy), a real /watch audit surface with SSE progress state, and behavioral fixes (fail-closed ranking without OpenRouter, epoch retry, safer git auth/object checks) with matching tests—core lasting product capability. B’s fail-closed vote auth, mock-OAuth gate, Secure cookies, and open-redirect hardening are high-value security fixes, but narrower in scope than A’s end-to-end deploy-and-audit system.
Side A delivers substantial new operational capability: production deployment (Docker, Fly.io, CI/CD), a live audit dashboard with SSE-backed progress reporting, status APIs, authenticated GitHub access for repository discovery, improved startup/retry behavior, and tests covering the new functionality. Side B contains valuable security hardening—failing closed on unauthenticated voting, Secure cookies, tighter redirect validation, pinned git dependencies, and small UX improvements—but it is a narrower refinement of existing authentication rather than adding comparable long-term platform capability.
comparison · c_7a129e904906 (tommy-mor) vs c_e2ee16c7ada5 (tommy-mor)
Side B is a genuine architectural cleanup: it separates identity/canonicalization concerns into dedicated modules, fixes a real inconsistency (identity rewriting scattered across reducer/events), removes dead double-@ sigil confusion, and makes the delegate optional with proper validation, backed by updated tests across CLI, server, and integration suites. Side A is mostly infrastructure/deploy scaffolding (Dockerfile, fly.toml, CI) plus a large CSS/JS dashboard UI bolted onto one file, which is useful operationally but adds more incidental complexity and less durable architectural value than B's identity refactor.
A ships production deploy (Docker/Fly/CI), multi-repo roots, and a real audit/SSE/status/watch path with tests and emission-loop hardening—net new lasting capability. B is a worthwhile identity/path split and wire-contract cleanup (optional naked delegate, no reducer rewriting), but mostly reorganizes and ripples existing behavior rather than adding comparable operational surface.
Side B makes a lasting architectural change by separating path normalization (`canonical_path.rs`) from identity parsing (`identity.rs`), removing identity rewriting from the reducer, enforcing strict stored-form usernames/agent IDs, and updating APIs so wire formats are consistent while HTML alone adds `@`/`@@` for display. Side A adds valuable deployment infrastructure, production configuration, and a live audit dashboard with SSE/status endpoints, but much of its patch is operational/UI work rather than the core data-model and API correctness improvements introduced in Side B.
comparison · c_7a129e904906 (tommy-mor) vs c_25172cf8caa0 (tommy-mor)
Side B fixes a real concurrency bug (shared sentinel delegate causing 'delegate already bound' failures that silently blocked every human user after the first vote), with a targeted type-level fix (Option<String> for agent), removal of now-dead sentinel logic, and a concrete regression test proving multi-user voting works. Side A is a large infra/UI deployment commit (Dockerfile, fly.toml, dashboard, SSE audit feed) that adds real operational value but is mostly additive scaffolding and cosmetic dashboarding rather than fixing a functional defect.
A ships lasting production infrastructure (Dockerfile, fly.toml, main-branch test-and-deploy CI), wires real repo/contributor roots and GitHub auth, and adds an auditable /watch SSE progress API with tests and epoch failure retry—operationalizing the constitution system end-to-end. B is a precise, high-leverage correctness fix (optional agent, drop shared WEB_BROWSER_AGENT so multi-user browser votes no longer hit AgentBound), but it is narrower in scope than A’s deployable runtime and observability surface.
Side B fixes a concrete correctness issue by removing shared browser sentinel delegates, changing pending sessions to use an optional delegate for browser logins, and ensuring browser-authored actions no longer create incorrect agent bindings that blocked multiple human users. Side A adds valuable deployment infrastructure, monitoring UI, SSE audit streaming, and production configuration, but it also bundles a large amount of operational and UI work whose long-term value is broader rather than addressing a core correctness bug; B's focused behavioral fix is more fundamental and is backed by integration tests covering the multi-user voting scenario.
comparison · c_7a129e904906 (tommy-mor) vs c_0a9a8eab32ba (tommy-mor)
Side B fixes a genuine correctness bug (feed catch-up relying on timestamps that can collide or move backwards) with a principled ingest-order anchor, deduplicates duplicated code paths into a shared rpc_feed helper, and adds thorough targeted integration tests for multi-user/private-room visibility. Side A is largely infrastructure/deploy plumbing (Dockerfile, fly.toml, CI) plus a UI/dashboard feature with generated CSS/JS, which is useful operationally but lower in durable design value and correctness impact than B's bugfix.
A lands durable production infrastructure (Dockerfile, fly.toml, gated deploy workflow), epoch retry/error handling, GitHub-authenticated discovery, and a tested live audit SSE/watch surface that makes the constitution process operable. B is a precise, well-tested feed fix (ingest-index anchors, private-room visibility, room on FeedPost) but is a narrower correctness patch than A’s end-to-end production and observability delivery.
Side B fixes a core correctness issue in feed retrieval by anchoring implicit catch-up to durable ingest order instead of timestamps, preserving correct behavior across same-millisecond posts, clock rollback, and permission changes. It also refactors the feed logic into a shared helper, includes room metadata, and adds thorough integration tests covering multi-user visibility and revoked access, whereas Side A is largely deployment, monitoring, and operational infrastructure with only a few smaller correctness improvements (such as deduplicating Git object verification and authenticated Git fetches).