Side B adds a new end-to-end browser test that seeds a pool, exercises the vote→next-pair workflow, and verifies both edge-history updates and that every displayed pair stays within the requested pool, providing lasting regression coverage for an important user flow. Side A mainly removes now-unused parameters and stops updating the vote-compare preview, simplifying the implementation but largely acting as cleanup with comparatively limited functional impact.
constitution · epochs · watch · epoch 3 · comparison · attempt
judgment
openai/gpt-chat-latest → B (4:1)
jud_f3428329d1c550 · raw event
Metadata
judgment_idjud_f3428329d1c550ddfdd05b4be9d6e83dd51878fa1a62b924cc50a64ce9e3ba31
model_idopenai/gpt-chat-latest
winnerB
ratio4:1
comparison_idcmp_2268e0cdd548c54a67e1c6bc1a2cb893591825604c3ed395423a33d28286078f
attempt_idatt_72638fe4e5d616d50654716871b19b8d6e5668938f8e6533d8d0ffd2750c370b