Media · AI & automation
Poder en los MediosA news portal with an AI pipeline that reads the press, groups each story and writes it up
News portal from Ecuador with an AI scanner that reads 20 outlets, groups the same story by semantic similarity and writes it up as an original piece.
Technologies
- Python
- FastAPI
- SQLAlchemy
- PostgreSQL
- Qdrant
- OpenAI API
- newspaper3k
- BeautifulSoup
- Next.js 16
- React 19
- TypeScript
- Supabase
- Tailwind CSS v4
- node-cron
- Zod
- Docker
- Railway
- Google Indexing API
- IndexNow

20
outlets monitored
3 min
between scanner sweeps
3,072
dimensions per embedding
18
endpoints in the scanner API
1,940
articles in Actualidad
3,636
URLs in the sitemap
My role
Full stack developer at Azirgo SAS, working alongside another developer who started the portal and set up the scanner's embeddings, job queue and similarity search. I built the per-outlet extractors, the AI grouping and writing, the newsroom monitoring panel, the portal integration (sync, auto-publishing, images and indexing) and the AI tools in the editor.
The problem
A small news portal competes with entire newsrooms. Covering the day means reading dozens of front pages, noticing that five outlets are telling the same story, writing your own version and publishing it before the news goes cold, without copying anyone and without losing the credit on every photo. Doing it by hand does not keep up, and automating it without control fills the site with duplicates, copied text and images hot-linked from other people's servers.
What I built
- A Python scanner that sweeps the front pages and sections of 20 outlets in continuous cycles with a 3-minute pause between them, among them El Universo, Primicias, El Comercio, Ecuavisa, Expreso, Extra, El País, BBC Mundo, CNN en Español and Infobae, with a dedicated extractor for each.
- Extraction of title, body, date, source section, caption and photo credit, with a credit detector that recognises agencies (EFE, AFP, Reuters, Getty…) and institutions however the caption is written.
- OpenAI embeddings for every story indexed in Qdrant, plus a table with each article's closest neighbours and their scores.
- Automatic grouping: stories from different outlets about the same event are gathered into a single group, and the ones that arrive late join the group that was already published.
- AI writing of the most complete version in each group: an original HTML piece with an SEO title, standfirst, seven paragraphs and three subheadings, with no invented facts.
- A FastAPI REST API that exposes the groups ready to publish, with a cursor, a pending-only filter and delivery confirmation.
- Sync in the Next.js portal: a cron job that pulls pending groups, re-hosts the image on Supabase Storage, creates the article and notifies search engines.
- An auto-publishing rule: if the photo is credited to a news agency or a public institution, the story goes live; otherwise it stays as a draft for the newsroom.
- Replicador de Noticias: a live panel for the newsroom with the feed per outlet, a spoken alert for every new story, semantic search and the same story seen across several outlets.
- AI tools in the portal editor: a title checked against the web, an SEO standfirst, internal-link suggestions and a generated editorial image.
How it's structured
Three independent projects, each with its own repository and its own Docker deployment on Railway, talking to each other over HTTP.
- news_scanner (Python 3.12, FastAPI, SQLAlchemy): a single process with several threads. The main one runs the three sweeps, news, sports and World Cup, back to back with a 3-minute pause between cycles; two threads (configurable) generate embeddings; two more group and write, news every 25 minutes and sports every 15; and Uvicorn serves the API. Data lives in PostgreSQL, across 8 tables, and vectors in Qdrant.
- Poder en los Medios (a Turborepo monorepo): the Next.js 16 portal on Supabase (PostgreSQL, Auth, Storage and RLS), the admin panel and the cron job that consumes the scanner, which starts inside the Next.js server itself. A separate FastAPI service feeds the site's job board.
- Replicador de Noticias (Next.js 16, React 19): the newsroom's internal panel. It has no database of its own: it reads the scanner API from the browser and uses two server routes as proxies, one for semantic search and one for downloading images.
- The scanner has two outputs: general-news groups, consumed by this portal, and sports groups, consumed by another portal in the group through the same mechanism.
How the pieces talk to each other
The whole flow is pull over REST: nothing pushes data to the portal, the portal asks for it and acknowledges it. State travels in two markers, sent_to_cms in the scanner and external_id in the portal, so any step can be repeated without duplicating anything.
- Sweep: each extractor reads up to 20 links from the outlet's front page, skips URLs already stored and stops at the first repeated headline. Every new story is saved with its category and a SHA-256 hash of title and body; if the hash is new, an embedding.request event is queued in the outbox table.
- Indexing: workers claim the event with FOR UPDATE SKIP LOCKED, request the embedding, look up the 15 nearest neighbours in Qdrant, store the pairs in similar_articles and record the vector. Up to 3 attempts per event; if Qdrant does not answer, the event goes back on the queue without using up an attempt.
- Grouping: every 25 minutes it takes the already-indexed stories from the 6 enabled Ecuadorian outlets, at least 1,000 characters long and 30 minutes old. They are grouped at a similarity of 0.72 or higher, the most complete one is written up and the group lands in published_news_articles with sent_to_cms empty.
- Consumption: the portal calls GET /published-news?unsent_only=true&limit=50 twice an hour, on the hour and at 45 past, between 09:00 and 17:45, and on the hour the rest of the day, Ecuador time, with a guard that stops two runs from overlapping. The response is validated with Zod before anything touches the database.
- Creation: a unique slug, the image downloaded and normalised to 1200×900 with WebP derivatives in Supabase Storage, an INSERT into articles with external_id, then POST /published-news/{id}/mark-sent.
- Extended groups: if an outlet joins a group after it was delivered, the scanner clears sent_to_cms; the portal receives it again, hits the unique index on external_id and only updates the source list, leaving title, body and slug untouched.
- Publishing: auto-published stories notify Google's Indexing API, IndexNow and the feed's WebSub hub in parallel. Drafts wait for an editor, who publishes from the panel and triggers the same notice.
- Live newsroom: the Replicador polls /articles/latest every 30 seconds and, when the id changes, shows a notification and reads the headline aloud with the browser's Web Speech API.
AI in the pipeline
AI does different jobs with different models, and none of them decides on its own what gets published.
- Embeddings: OpenAI's text-embedding-3-large, 3,072 dimensions, over the title and body of each story. The output is a vector in Qdrant with cosine distance plus the similarity pairs. It is what makes it possible to recognise the same story told in different words by another outlet, something no headline comparison can do.
- Semantic search: the same model turns what the newsroom types into the Replicador into a vector, and Qdrant returns the closest articles from the last few hours with their similarity percentage.
- News writing: a configurable GPT model, gpt-5.4-nano by default, through Chat Completions. It receives the title, standfirst and body of the most complete article in the group and returns JSON with a title, a standfirst and an HTML body of 7 paragraphs and 3 subheadings, with no invented facts and no traces of live coverage or tweets. If the lengths are off, the error is sent back so it can correct itself, with 3 attempts in total.
- Sports: the same model with a sports-editor prompt that turns minute-by-minute coverage into post-match stories, for the scanner's second output.
- Title in the editor: OpenAI's Responses API with a versioned prompt that searches the web. It returns a proposal and alternatives; if it cannot confirm the facts it answers NO_CONFIRMADO and the headline stays as it was instead of being filled with something invented.
- SEO standfirst: another versioned prompt that takes the title and the current standfirst and returns a new standfirst along with alternatives.
- Internal links: gpt-3.5-turbo suggests phrases in the body that deserve a link; the portal searches its own archive for a story that covers each one and links only those that find a target.
- Image: gpt-image-1-mini generates a text-free editorial image from the headline, the standfirst and the editor's instructions, or edits one from a reference photo, at 1536×1024 for the featured image or 1024×1024 for Instagram, and uploads it straight to Supabase Storage.
Inside the site

Diagram: how the scanner, the AI, the portal and the Replicador talk to each other. 
Actualidad, the section the scanner pipeline feeds. 
An Actualidad story, with headline and standfirst within the pipeline limits. 
The story body with the structure the AI rewrite asks for: paragraphs and subheadings. 
Replicador de Noticias, the newsroom panel (running locally on sample data). 
Live feed per outlet, with source section and photo credit (sample data). 
The same story across several outlets, ranked by semantic similarity (sample data).
The admin panel
Each module covers one part of the day-to-day operation, without leaving the panel.
Per-outlet extractors
One module per outlet that reads its front page or section, skips what is already stored and stops at the first repeated headline.
Semantic similarity
Embeddings in Qdrant plus a table of scored pairs used by grouping, the panel and search.
Grouping
Gathers the same story told by several outlets and adds late-arriving sources to a group already published.
AI writing
An original HTML piece with title, standfirst, seven paragraphs and three subheadings, validated before it is saved.
Sync
A portal cron job that pulls pending groups, creates the article and confirms delivery back to the scanner.
Auto-publishing
A story with an agency or public-institution photo goes live on its own; the rest stays as a draft.
Instant indexing
Every publication notifies Google's Indexing API, IndexNow and the feed's WebSub hub in parallel.
Replicador de Noticias
The newsroom's live panel: feed per outlet, spoken alerts, semantic search and related stories.
Technical decisions
- A PostgreSQL outbox between the scraper and the embeddings: the scraper never waits on OpenAI, and if Qdrant or the API goes down the work goes back on the queue instead of being lost. Workers claim events with SKIP LOCKED so they never step on each other.
- Similarity is computed once, when each story is indexed, and stored as scored pairs. Grouping, the newsroom panel and the scanner's second output read that table instead of querying the vector index again.
- Only the most complete version of each group, the one with the longest body, gets written up: one model call per story, not one per outlet.
- The model answers in JSON with hard limits, a 45 to 70 character title and a 120 to 320 character standfirst. If it goes over, the error is sent back so it can correct itself, and as a last resort the standfirst is cut at a full sentence.
- A pull integration with acknowledgement: the portal asks for what is pending and marks each group as sent. A unique index on external_id means a retry can never duplicate a story.
- What gets published without review is decided by a readable, dependency-free rule, the photo credit, not by the model. Everything else goes through an editor.
- Images are downloaded and re-hosted after checking the file's real bytes, because many CDNs answer with an HTML page and a 200 status when they block hot-linking.
Challenges
- Every outlet builds its HTML its own way: it took one extractor per outlet, and fixing cases like a site that did not declare its encoding in the header and left broken accents across thousands of articles.
- Tuning the similarity threshold so two stories about the same event group together without merging different stories, and waiting 30 minutes before grouping so the other outlets have time to publish theirs.
- Letting an outlet that publishes hours later join a group already delivered without overwriting the title or the text the newsroom is editing in the portal.
- Reading the photo credit from captions written in very different ways, «Foto: EFE», «Foto AFP» or the agency alone at the end, without mistaking the rest of the sentence for a name, because that credit decides what publishes on its own.
Result
The portal keeps its Actualidad section current with stories the pipeline finds, groups and writes up. Agency and official-source stories go live on their own, with the image hosted on the site itself and an immediate notice to Google, IndexNow and WebSub; the rest reach the editor as drafts already written and with their sources identified. The newsroom works from a live panel that alerts it to every new story and shows the same story across several outlets, instead of checking front pages one by one.