All projects
Data Engineering · Quant MLFEATURED

Financial News Market Prediction Pipeline

100 million financial news articles scraped across 20 servers and filtered down to 2 million, then run through everything from a fine-tuned FinBERT to attention-based time series networks to try to call where a stock was going.

→ 100M articles scraped across 20 servers, filtered to 2M relevant
Source private, exploring this as a startupStock Prediction ReportOil Price Analysis
PythonFinBERTDoc2VecCatBoostBiLSTMTransformersPlaywrightGoLoginBigQuery

// THE STORY

This started as a question: if you could read every financial news article published about a company over the past few weeks, could you predict where its stock was going the next day? My team and I spent a long time on that, and it pulled us much deeper into data engineering, modeling, and reasoning text model orchestration than any of us expected going in.

The first real problem was collection. We built a scraping pipeline across 20 remote JHU servers that gathered roughly 100 million financial articles, pulling article URLs from Google News and BigQuery's Global News Knowledge Graph and stripping the ads and boilerplate out of the raw HTML.

Almost none of that volume is useful, because an article that mentions a ticker is not necessarily an article about the company. We embedded each document with Doc2Vec and scored it against a text description of the company it was supposed to be about, which cut 100 million articles down to the 2 million that were actually relevant.

We then fine-tuned the later layers of a 250 million parameter FinBERT and several other language models, replacing their final layers with our own N-way classification so they predicted straight from the article text. In parallel we engineered features out of the news and fed those to models that want structured input instead, ensembling CatBoost and recurrent LSTMs, so the text models had something to be measured against on even footing.

Then we moved toward higher frequency prediction, which put all the pressure back on the scraper. Intraday signal needs articles bounded by exact timestamps, so we built a much more capable invisible scraper that could retrieve information on the backs of powerful search engines for a specific window without getting blocked or flagged, and return what was published between two times rather than whatever happened to be indexed later.

We benchmarked the forecasting side, from decision trees up to attention-based time series networks, across returns for 10 tech stocks. Afterward I pointed the same scraper at crude oil coverage during the US-Iran conflict, on the theory that a sharper and more news-driven market might outperform intuition.

The results we got for those models were high variance and barely beat the S&P 500 or random choice. What worked better was agentic orchestration, where different agents take the role of different stock market analyzers and debate how the classification is going to turn out. That is currently being explored as a startup.

// WHAT STANDS OUT

Mostly throwing data away

A ticker mention is not coverage. 98 of every 100 articles got discarded.

Publish time, not index time

Search engines tell you when they found an article. Intraday needs when it went up.

Barely beat a coin flip

FinBERT, CatBoost ensembles, attention time series nets. High variance, barely past random choice.

Making them argue worked

Agents in different analyst roles, debating the call, beat every model we fine-tuned.

This came out of a role on my timeline. See the experience entry