Skip to content
HuzzFlow

HuzzFlow Insights

How to spot a machine-generated news article

Synthetic articles are fluent, confident, and specifically wrong. Fluency is no longer evidence of effort.

HuzzFlow Editorial Desk 4 min read

Updated

Stop using writing quality as a filter

For most of the web's history, clean prose was a reasonable proxy for effort. Someone competent had spent time. That heuristic is now obsolete: fluent, well-organized, grammatically flawless text is cheap, and a large volume of low-value content reads better than it used to.

This does not mean machine assistance in journalism is inherently illegitimate — plenty of newsrooms use it for transcription, translation, and first drafts under human editing. The problem is the fully automated article published without anyone checking whether it is true. Detecting that is now a basic reading skill, and it depends on different signals than style.

Specificity is where synthetic text breaks

Generated articles are excellent at the shape of a news story and unreliable at its verifiable particulars. Look closely at the details a human reporter would have had to obtain from somewhere: the exact title of an official, the precise figure and its units, the date of a filing, the name of the court, the quarter a result belongs to.

You will often find these are plausible and wrong — the official held that role two years ago, the figure is real but from a different report, the company name is subtly incorrect. The article reads authoritatively because the language is confident, and the errors cluster exactly where reporting effort would have been required.

Check every quote

Quotes are the most reliable test available to a reader. A real quote came from a person, in a setting, at a time — a press conference, an interview, an earnings call, a written statement. It is therefore findable elsewhere, or at minimum attributable to an occasion.

Search a distinctive fragment of the quote in quotation marks. Genuine quotes from public events appear in multiple places or in a transcript. Fabricated ones appear only in the article you are reading, and often only on that domain. A story built on quotes that exist nowhere else, attributed to people who are hard to locate, is not reporting.

Be similarly wary of the generic expert: a named analyst with an institution you cannot find, saying something no specific person needed to say. Real sourcing is particular, because obtaining it required reaching a particular person.

Read for institutional traces

Actual journalism leaves fingerprints of the organization that produced it. A named reporter with a beat and an archive. An editor's note when something changes. A dateline naming where the reporting happened. Links to the filing or the ruling. A photograph with a credit. House style applied consistently — how the outlet renders numbers, honorifics, and dates.

Automated content typically has none of this. No byline or a house byline, no dateline, no outbound links to primary documents, stock imagery with no credit, and no consistent style because there is no style guide. Any one absence is unremarkable. All of them together describe a page with no newsroom behind it.

Look at the site, not just the article

The strongest signals are often visible one level up. Open the front page and check publication volume against apparent staff size — dozens of articles daily from a site listing two people is arithmetic that does not work. Check topical coherence: automated operations chase whatever has search volume, so unrelated high-traffic subjects sit next to each other with no editorial logic.

Then read three articles in a row and notice structural sameness. Generated content tends toward identical architecture: same section count, same paragraph rhythm, same hedged conclusion recommending that readers stay informed and consult professionals. Human writing under deadline is messier and more variable than that.

Beware the plausible middle

The genuinely difficult case is not obvious junk. It is an article about a real event, largely accurate, with two or three fabricated specifics woven in — a quote nobody said, a number from the wrong year, a causal claim no source made.

Nothing about the surface warns you, and the surrounding accuracy lends the invented parts credibility. The only defence is the habit from the previous sections: verify the particular detail you intend to rely on, rather than judging the article as a whole. Assess claims individually, at the granularity you plan to use them.

A working checklist

Applied to anything you are considering acting on or resharing:

  • Is there a named author with a traceable archive and beat?
  • Do the quotes appear anywhere else, or in a transcript of a real occasion?
  • Are the named experts and their institutions findable?
  • Does the article link to the primary document it describes?
  • Do one or two checkable specifics — title, figure, date, units — hold up?
  • Does the site's publishing volume make sense for its stated staff?
  • Do three articles from this site share suspiciously identical structure?

Where this leaves the reader

The practical adjustment is a shift in what earns trust. Not polish, which is now free, but traceability: named people, findable sources, linked documents, and a visible record of being corrected. Those are expensive to fake at scale precisely because they connect a page to accountable humans.

That shift also raises the value of a short, deliberately maintained list of sources you have actually vetted. In an environment where anyone can generate a convincing newspaper in an afternoon, knowing whose work you trust — and why, on which beat — is the durable skill.