Fix duplicate feed items from concurrent syncs

A sync button press racing a page-reload sync could run two sync
requests concurrently. create_feed_item used a check-then-insert
pattern (query for an existing item, then insert if none found), so
both requests could pass the check before either inserted, creating
duplicate items.

- Add a UNIQUE (feed_id, url) constraint on feed_item, keyed on the
  article's link rather than its title since that's RSS's stable
  article identity (a feed editing a headline after publishing, same
  link, would otherwise dupe under title-based matching).
- The migration first collapses any duplicates already created by the
  race, propagating read=true onto the surviving row when any
  duplicate in its group was already read, so the cleanup can't make
  an already-read article look unread.
- create_feed_item now does a single atomic
  INSERT ... ON CONFLICT (feed_id, url) DO NOTHING instead of
  select-then-insert, closing the race entirely.
- Add a test that races two threads on separate connections inserting
  the same item and asserts only one row survives.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012CHcxDSPbHhe7sLJQBRVC9
This commit is contained in:
2026-09-10 16:04:34 +02:00
co-authored by Claude Sonnet 5
parent e2fb32f898
commit 6c8b4e9e32
3 changed files with 115 additions and 19 deletions
@@ -0,0 +1,3 @@
-- This file should undo anything in `up.sql`
ALTER TABLE feed_item
DROP CONSTRAINT feed_item_feed_id_url_unique;
@@ -0,0 +1,36 @@
-- Your SQL goes here
-- Concurrent syncs for the same feed could previously race past the
-- application-level "does this article already exist" check and both
-- insert, producing duplicate feed items. Before enforcing uniqueness,
-- collapse any duplicates that already exist, keyed on the article's link
-- (rather than its title) since that's the stable identity RSS items are
-- built around — some feeds edit a headline after publishing while keeping
-- the same link, which title-based matching would treat as a new article.
-- If any duplicate in a group was already read, propagate that onto the
-- surviving (lowest-id) row first, so this one-time cleanup can't make an
-- already-read article look unread again.
UPDATE feed_item AS keep
SET read = true
FROM feed_item AS dupe
WHERE keep.feed_id = dupe.feed_id
AND keep.url = dupe.url
AND keep.id <> dupe.id
AND keep.id = (
SELECT MIN(candidate.id)
FROM feed_item AS candidate
WHERE candidate.feed_id = dupe.feed_id
AND candidate.url = dupe.url
)
AND dupe.read = true;
-- Keep the oldest (lowest id) row of each duplicate set.
DELETE FROM feed_item a
USING feed_item b
WHERE a.feed_id = b.feed_id
AND a.url = b.url
AND a.id > b.id;
ALTER TABLE feed_item
ADD CONSTRAINT feed_item_feed_id_url_unique UNIQUE (feed_id, url);