Written by Kartik Chugh
In Q3 2025 we came close to killing a video format for a client because it was our lowest performer on every metric we reported. Average watch time was poor, completion rate was worse, and across 6 weeks it had produced a fraction of the views of the formats we were proud of. We had the recommendation written. Then someone on the client’s sales team mentioned, in passing on a call, that prospects kept referencing it.
That sentence cost us a week of re-analysis and changed how we report on video across 9 clients.
What we were measuring
The format was a plain screen recording, 4 to 6 minutes, of someone using the product to do one specific task. No script, no editing, no music. We produced them because the client’s team liked them, and we quietly considered them a distraction from the work that generated reach.
Our reporting was standard. We pulled views, average watch time and completion from the platform, and we ranked formats by those numbers in a monthly summary. The short, sharply edited pieces won every month. The task recordings lost every month, consistently and by a wide margin.
Every one of those numbers was accurate. The ranking built on them was close to meaningless, and I could not have told you why at the time.
What was actually happening
The two formats were doing entirely different jobs and we were scoring them on the same scale.
The short pieces created awareness. They reached people who did not know the product existed, most of whom were never going to buy, which is exactly what awareness content is supposed to do. High view counts, shallow engagement, and that is fine.
The task recordings were doing something else. Almost nobody watched them. The people who did were late in a buying process, trying to answer a specific question about whether the product actually did the thing. They watched with intent, they watched more than once, and a meaningful share of them then talked to sales already knowing what they wanted.
Our metrics were built to measure the first job. Applied to the second, they described a failure. A video watched by 200 people who are choosing a vendor is not underperforming a video watched by 20,000 people who are scrolling, and our monthly ranking said it was.
How we sized it
We did the only thing that settles this, which is asking the people who were on the calls. Over 4 weeks the client’s sales team logged, in HubSpot, whether a prospect referenced any specific content during a first call. It was a single field and it took them seconds.
The result was not subtle. The task recordings appeared in roughly a third of logged first calls. The high-reach pieces appeared in almost none. When we compared that against which conversations progressed, the pattern held.
We had been about to delete the thing that was doing the work, on the strength of a report I had built and defended.
In retrospect the signal had been available for months and we had filed it as noise. The client’s team had mentioned twice in quarterly reviews that customers brought these videos up. Both times I heard it as a preference rather than as data, because it arrived as an anecdote and our reporting arrived as a table. The bug was not in the metric. It was that we had an implicit rule, never written down, that numbers outrank observations, and that rule is wrong whenever the numbers cannot see the event being observed.
What we changed
First, we stopped ranking formats against each other. Every format is now tagged in our reporting as either reach or consideration, and it is only compared to other formats doing the same job. That change alone eliminated most of the bad recommendations we were generating, because the vast majority came from comparing across the boundary.
Second, we added the sales-mention field permanently. It is the cheapest instrument we have for anything that happens off-platform, and it is the only one that sees the part of the funnel our analytics cannot. It is self-reported and imperfect and still better than the alternative, which was inferring intent from watch time.
Third, we changed what we produce. The client now makes more task recordings, not fewer, and we stopped putting production effort into them because the roughness is part of why they are trusted. A polished version of that video would be an advert. The unpolished one reads as evidence.
We have since run the same tagging across 9 clients over 7 months. The split holds in 7 of them. The 2 where it does not are both companies selling to a buyer who does no self-directed research before talking to sales, which is itself a useful thing to learn about an account early.
What I would tell someone running video
The failure was not a bad metric. Watch time is a fine metric. The failure was applying one scoring system across content doing structurally different jobs, and then acting on the ranking that it produced.
If you take one thing, take the sales-mention field. Ask your sales or support team to log whether a prospect referenced anything specific, for a month. It costs almost nothing and it will tell you things your analytics is structurally unable to see, because the moment that matters happens in a conversation your platform was never present for.
The second thing I would say is to be suspicious of any format your own team likes that your reporting says is failing. That combination is not usually sentimentality. It is often a sign that the people closest to customers are seeing something your instruments are not built to capture. We nearly overruled exactly that signal because we had numbers and they had an impression, and the numbers were answering a different question. How we think about video that carries weight is in our clipping campaign mistakes piece.
The recommendation to kill the format is still in a draft deck somewhere. I keep it as a reminder that a confident, well-evidenced, internally consistent recommendation can be completely wrong if the measurement underneath it was built for a different purpose.
Author Bio:
Kartik Chugh (Simba) is a founder-operator at the intersection of distribution, culture, and narrative control in Web3.
Cofounder of FORKOFF, a culture and distribution studio that designs IP-driven campaigns, event systems, and narrative loops for protocols, funds, and builder ecosystems. FORKOFF treats events as content factories, founders as distribution engines, and culture as infrastructure — not aesthetics. 3,085+ short-form clips every 13 days for clients. $5M+ in ecosystem activations across 14 countries.
Previously CMO at QuillAudits, the Web3 security pioneer, where he scaled security products to 100K+ users, built 150+ ecosystem partnerships, generated $3M+ qualified pipeline, and drove 1Bn+ views across campaigns. Co-founded EdSquare (acquired). Five years across the AI, Web3, and B2B SaaS playbook.
Hosted and partnered on 100+ global events across ETHDenver, Token2049, Consensus, Devcon, and KBW in 20+ countries. Leads Misfits Dubai, a founder-first community built around curated rooms rather than mass communities. Builder at Seedrail (the distribution stack for tech and VCs). Active investor in 12+ early-stage startups across crypto and AI.
Frequent contributor to CoinDesk, CoinTelegraph, The Defiant, and Block Telegraph. Speaker at Token2049 Singapore and QuillCon. Advisor at TiE Global and ADSME HUB.
Speaks on: founder-led distribution, events as content factories, rooms > reach, culture > campaigns, narrative control in Web3, and creator-led distribution.
Available for commentary on: AI agency growth, Web3 marketing, podcast clipping ROI, founder-led GTM, KOL marketing, and ecosystem activation strategy. Based in Dubai.