Daily News & Article Data Collection from 24 Sites
The Client A US organisation running a media-intelligence function — the kind of operation that needs to know, every morning, what was published overnight across a defined set of sources that matter to its domain. Their source list was specific and mixed: 24 websites spanning mainstream news outlets, trade and industry publications, and community forums. Their need was equally specific: a clean, daily, structured feed of everything newly published, delivered in a consistent shape they could analyse, search, and route — regardless of how differently those 24 sites were built. The brief was, on its surface, one of the most common data requests in existence: "collect the articles from these sites, every day." Its difficulty — and the reason it's worth documenting — is that "these sites" spanned 24 completely different content architectures, and "clean" turned out to be doing enormous work in that sentence. The Challenge News and article extraction is de...