Fix duplicate feed items from concurrent syncs
A sync button press racing a page-reload sync could run two sync requests concurrently. create_feed_item used a check-then-insert pattern (query for an existing item, then insert if none found), so both requests could pass the check before either inserted, creating duplicate items. - Add a UNIQUE (feed_id, url) constraint on feed_item, keyed on the article's link rather than its title since that's RSS's stable article identity (a feed editing a headline after publishing, same link, would otherwise dupe under title-based matching). - The migration first collapses any duplicates already created by the race, propagating read=true onto the surviving row when any duplicate in its group was already read, so the cleanup can't make an already-read article look unread. - create_feed_item now does a single atomic INSERT ... ON CONFLICT (feed_id, url) DO NOTHING instead of select-then-insert, closing the race entirely. - Add a test that races two threads on separate connections inserting the same item and asserts only one row survives. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012CHcxDSPbHhe7sLJQBRVC9
This commit is contained in:
@@ -0,0 +1,3 @@
|
||||
-- This file should undo anything in `up.sql`
|
||||
ALTER TABLE feed_item
|
||||
DROP CONSTRAINT feed_item_feed_id_url_unique;
|
||||
@@ -0,0 +1,36 @@
|
||||
-- Your SQL goes here
|
||||
|
||||
-- Concurrent syncs for the same feed could previously race past the
|
||||
-- application-level "does this article already exist" check and both
|
||||
-- insert, producing duplicate feed items. Before enforcing uniqueness,
|
||||
-- collapse any duplicates that already exist, keyed on the article's link
|
||||
-- (rather than its title) since that's the stable identity RSS items are
|
||||
-- built around — some feeds edit a headline after publishing while keeping
|
||||
-- the same link, which title-based matching would treat as a new article.
|
||||
|
||||
-- If any duplicate in a group was already read, propagate that onto the
|
||||
-- surviving (lowest-id) row first, so this one-time cleanup can't make an
|
||||
-- already-read article look unread again.
|
||||
UPDATE feed_item AS keep
|
||||
SET read = true
|
||||
FROM feed_item AS dupe
|
||||
WHERE keep.feed_id = dupe.feed_id
|
||||
AND keep.url = dupe.url
|
||||
AND keep.id <> dupe.id
|
||||
AND keep.id = (
|
||||
SELECT MIN(candidate.id)
|
||||
FROM feed_item AS candidate
|
||||
WHERE candidate.feed_id = dupe.feed_id
|
||||
AND candidate.url = dupe.url
|
||||
)
|
||||
AND dupe.read = true;
|
||||
|
||||
-- Keep the oldest (lowest id) row of each duplicate set.
|
||||
DELETE FROM feed_item a
|
||||
USING feed_item b
|
||||
WHERE a.feed_id = b.feed_id
|
||||
AND a.url = b.url
|
||||
AND a.id > b.id;
|
||||
|
||||
ALTER TABLE feed_item
|
||||
ADD CONSTRAINT feed_item_feed_id_url_unique UNIQUE (feed_id, url);
|
||||
Reference in New Issue
Block a user