This interview is with Dr Hesham Mashhour, Founder, HarperFlow.
Introduce yourself: as a founder in higher education with research and journalism roots, what problems are you focused on solving at the intersection of geo articles, source-backed content, and Webflow SEO?
I9m Hesham Mashhour, founder of HarperFlow. My background is a bit of a zigzag: I trained as a medical doctor at the University of Cambridge, then spent years making research readable for general audiences founding NeuroEverything, a neuroscience-education channel, writing for HuffPost UK, and teaching online courses. The common thread is taking source material seriously.
The problem I9m focused on: the way buyers find things changed faster than the way companies publish. People now ask ChatGPT, Perplexity, and Gemini what to buy. In research I ran across 117 AI answers, the engines barely agreed with each other. 79% of the 646 brands in the data existed in only one engine’s world. Most companies are invisible to these engines and don’t know it. In a separate index of 1,464 Webflow agencies, only 4.2% were ever named by any engine.
The pages AI engines do cite follow a pattern: source-backed, structured, specific. Pages with FAQ blocks and tables were cited 75% of the time, versus 22% for pages without. Only about 31% of cited pages contain any original data so verifiable substance is rare enough to stand out.
Producing that content properly is slow: real research, real sources, and checking every claim. Most Webflow teams can’t justify a research desk. That’s the gap HarperFlow fills: every article researched across 812 sources, fact-checked per claim, and published straight to the Webflow CMS. Content built to be cited.
Looking back, which experiences in adult education, scientific research, or publishing most shaped your approach to building trustworthy, high-performing content?
Probably three things.
Medical training, first. At Cambridge, a claim without a source is basically gossip. That wiring never left.
Then NeuroEverything. Years of turning neuroscience papers into videos people actually finish taught me that accuracy and readability aren’t in tension — the readable version just takes more work. Most bad content fails on effort, not talent.
And publishing — HuffPost UK and books — showed me what editors cut: anything they can’t verify and anything that smells like marketing.
The approach that came out of all that was: every claim traced to a source, written plainly, and nothing you’d have to defend later. That became HarperFlow’s whole model.
Building on that, can you walk us through your end-to-end workflow for producing a source-backed article—from idea to publish—highlighting where claim-level verification lives?
It’s pretty mechanical by design.
-
It starts with the site, not the keyword. HarperFlow profiles the site first — what it sells, who reads it, and what it can credibly claim — because an article that doesn’t fit the site won’t get cited anyway.
-
Research. Every article pulls from 8–12 real sources, and research takes about 13 minutes per article. Sources are stored, not skimmed — they become part of the article’s data.
-
Drafting happens per section, and this is where verification lives: not a review at the end, but a per-claim check inside the pipeline. Each factual claim is traced back to its source before the article moves on. Nothing ships with an untraceable claim.
-
Structure — FAQ blocks, tables, named author — because pages built that way get cited far more often. Then the article goes straight to the Webflow CMS.
-
If any step fails its check, the article goes back, not out.
Zooming into geo content, how do you scale coverage across cities or states in areas like healthcare, education, or civic services without creating doorway pages or losing local nuance?
The honest answer is that doorway pages are a sourcing failure, not a structural failure.
A page becomes a doorway page when the only thing that changes between versions is the city name. The fix isn’t clever templates — every page must earn its existence with local substance: local data, local rules, local providers. If you can’t find genuinely different source material for a city, you probably shouldn’t have a page for it.
That’s how we think about it at HarperFlow. Every article is researched against 8–12 sources, and for a geo page those sources are local — state regulations, city-level data, local institutions. The per-claim fact-check then enforces the nuance: a claim that’s true for Texas gets flagged on the California page.
It scales because the research scales, not because the template does. It’s slower per page than a find-and-replace mill — obviously — but those mills are exactly what Google and the AI engines have learned to ignore.
Staying with distribution, after auditing ChatGPT, Perplexity, and Gemini, what specific choices most increased the odds that your pages or original research are cited or recommended by these assistants?
The biggest one: structure. In a study I ran analyzing 117 AI answers, pages with FAQ blocks and tables were cited 75% of the time versus 22% without. Same information, different packaging: three times the rate.
Second: original data. Only about 31% of cited pages contain any original data. Publishing our own numbers — even small, careful ones — puts us in that minority.
Third: stop chasing any single engine. The three engines agreed on a brand only 56% of the time, and Reddit went from appearing in 87% of Perplexity’s answers to 18% in fifteen days on identical questions. If you’d optimized for Reddit in week one, you’d have been optimizing for a ghost by week three.
An honest caveat: these raise the odds, but nobody controls the outcome. I build content to be cited — anyone who promises citations is selling something.
On original research, how do you design a small-but-rigorous study that produces quotable data publishers trust and assistants surface?
Small is fine. Sloppy isn’t.
My study was 39 buyer-style queries run through ChatGPT, Perplexity, and Gemini — 117 answers.
What made it work:
- Real questions. Buyer queries people actually ask, not prompts designed to produce a result.
- Fixed questions, repeated. The most interesting finding — Reddit collapsing from 87% to 18% of Perplexity’s answers — only exists because I re-ran the identical questions fifteen days later. A snapshot would have missed the whole story.
- Count everything, not just the interesting bits. 646 brands, 328 cited domains. Boring counting is what makes the percentages defensible, and most of those domains were cited exactly once — a detail you only find by counting all of them.
The trust part is simpler than people think: publish the method alongside the numbers. Publishers and AI engines both favor data they can interrogate, and with only ~31% of cited pages containing original data at all, a documented small study already clears most of the field.
Shifting to demand, how do you define “earning demand” in SEO based on a campaign where original insights created demand rather than just capturing keywords?
My definition: Earning demand is when people look for you because of something you published — branded searches, direct visits, and journalists or founders coming to you. Capturing demand is intercepting searches that already existed. The first compounds; the second stops the moment you stop feeding it.
The campaign is our State of AI Citations research, plus a free GEO audit tool, published ahead of HarperFlow’s launch. The logic: since only ~31% of the pages AI engines cite contain original data, publishing real numbers about what engines actually cite creates searches that didn’t exist before — people look for the data itself, not just the topic.
Being pre-launch, I’d rather give you the mechanism than a fake curve: insight leads to citation, citation leads to branded search, and branded search leads to inbound. That’s the sequence I’m betting on.
Bringing it to implementation, what is your blueprint for structuring Webflow CMS collections and schema so large, source-backed geo libraries stay fast, interlinked, and technically clean as they scale?
I’m a solo founder, so the blueprint is boring on purpose.
Three collections do most of the work:
- Articles
- Locations
- Sources
Locations and sources are their own collections, linked to articles through multi-reference fields — so a source or a city gets edited once and updated everywhere, and interlinking is structural instead of manual.
FAQs and tables are structured fields, not blobs of rich text. FAQ schema is generated from the same fields the page renders from, so the markup can’t drift away from the content.
Speed: Webflow publishes static pages, so the usual sins are oversized images and third-party scripts.
- Keep templates lean.
- Paginate big listing pages.
- Don’t bolt on a JS framework to render what the CMS already renders.
And because every article’s claims are checked against its sources before publishing, the sources collection doubles as the audit trail — every page can show its working.
Finally, what single leading indicator tells you your source-backed geo content is earning demand before conversions catch up?
Being named at all.
When I re-run buyer queries through the engines, the first thing I look for is whether the brand exists in the answer — not position, not sentiment, just presence. In my data, 79% of 646 brands appeared in exactly one engine’s world, and only 4.2% of 1,464 Webflow agencies were ever named by any engine. Most brands fail before ranking is even a question.
So the indicator is this: re-run your target queries every couple of weeks and count how often you’re named. Individual citations flicker — Reddit went from 87% of Perplexity’s answers to 18% in fifteen days — but a rising presence trend across all three engines means the content is compounding. Conversions usually trail presence by a while.
The boring second choice is branded search volume. It lags the engines slightly, but it doesn’t lie.
Thanks for sharing your knowledge and expertise. Is there anything else you'd like to add?
One thing, maybe.
Most of what I’ve said comes back to a simple observation: the engines people now ask for recommendations barely agree with one another, and most brands are invisible in all of them. That’s fixable—and it’s fixable with unglamorous work: real sources, clean structure, original data.
If you want to see where your own site stands, the full research is at harperflow.io/geo-research, and there’s a free GEO audit tool on harperflow.io. HarperFlow launches September 1, 2026 — until then, the research and the tool are the product.
Thanks for having me.