Video SEO in the AI Search Era: How to Get Your Videos Cited by ChatGPT, Perplexity, and AI Overviews

Video SEO in the AI Search Era: How to Get Your Videos Cited by ChatGPT, Perplexity, and AI Overviews

Your View Count Has Almost Nothing to Do With Getting Cited

Otterly's YouTube AI Citation Study, published March 2, 2026 and built on more than 100 million citation instances collected over 30 days across six AI platforms, found that the correlation between a video's view count and how often AI engines cite it is -0.03. Likes: -0.02. Channel subscribers: -0.03. Statistically, those are zeroes. In the same dataset, 40.83% of the YouTube videos that AI engines cited had fewer than 1,000 views.

That is a strange and useful fact. Every ranking system creators have optimized for since 2011 rewards accumulated popularity. The systems now answering a growing share of search queries do not. They are reading for a usable answer, and they will lift it out of a 400-view video from a channel nobody has heard of if that video happens to say the thing clearly.

Across 10,000+ projects and 200M+ views for clients, we have watched the referral mix shift under our roster over the last eighteen months. This is the first structural change to video discovery since suggested feeds overtook search, and almost nobody's upload checklist has caught up with it.

1. Long-Form Is the Format That Gets Cited

The same study breaks citations down by format: 94% went to long-form videos, 5.7% to Shorts, and 0.3% to playlists, channels, and livestreams combined.

AI engines overwhelmingly cite long-form video: 94% of YouTube citations went to standard uploads versus 5.7% to Shorts.
AI engines overwhelmingly cite long-form video: 94% of YouTube citations went to standard uploads versus 5.7% to Shorts.

The mechanism is not mysterious. A 45-second Short contains one claim and no supporting structure. A 14-minute explainer contains a dozen claims, each surrounded by the qualifying language a retrieval system needs to decide the passage is relevant. Length is not the point — extractable, self-contained passages are the point, and long-form is simply where they live.

This does not mean stop making Shorts. It means stop expecting Shorts to do a job they are structurally unsuited for. Shorts are a discovery and audience-building surface; that trade-off is laid out in our cross-platform short-form playbook. Citation is a long-form game.

2. Chapters Are the Unit AI Engines Actually Cite

31% of the cited videos in the study contained timestamp signals — and among those, 78% were cited multiple times across two to five different chapters of the same video. Google AI Overviews accounted for 73% of all timestamped citations and Google AI Mode for the rest; ChatGPT, Copilot, Gemini, and Perplexity produced none during the study window.

Read that again, because it changes how you structure a video. A well-chaptered 20-minute video is not one citable document. It is five. Each chapter is independently retrievable, independently surfaceable, and independently linkable with a ?t= deep link that drops the viewer at the exact second the answer starts.

Three rules we now apply to every long-form edit:

  • Chapter titles are queries, not labels. 04:12 — How much does a video editor cost? gets retrieved. 04:12 — Pricing does not. The chapter title is a heading in a document, and it should read like one.
  • Every chapter opens with its own answer. The first sentence after a chapter marker states the conclusion. Then it justifies it. A retrieval system that lands mid-video has no preceding context to work with.
  • Six to ten chapters on a 12–20 minute video. Fewer and each chapter covers too much ground to be a clean extract; more and they fragment below the length a model can use.

YouTube's own video chapters requirements set the floor: the first timestamp must be 00:00, there must be at least three of them in ascending order, and no chapter can be shorter than 10 seconds. Miss any one of those and the chapters silently don't render — which means no deep links, and nothing for a retrieval system to address.

3. The Transcript Is the Only Part of Your Video a Model Can Read

No AI search engine watches your video. It reads the transcript, the description, the title, and the metadata — nothing else. Your grade, your sound design, your cut are all invisible to it.

What actually gets cited: the transcript supplies the language, chapters give it structure, timestamps make a passage addressable, and the AI answer links back to that exact second.
What actually gets cited: the transcript supplies the language, chapters give it structure, timestamps make a passage addressable, and the AI answer links back to that exact second.

Which makes the caption file the single highest-leverage asset on the upload, and auto-generated captions a liability. YouTube's auto-captions mangle exactly the words that make a video worth citing: product names, technical terms, numbers, proper nouns. An uploaded .srt fixes all of it, and the reasoning behind that choice is in our guide to burned-in vs. closed captions. Burned-in text does nothing here — pixels are not text.

Two habits that follow from this, and both are scripting decisions rather than editing ones:

  1. Say the full noun before the pronoun. "The Sony FX3 shoots 4K120" is extractable. "It shoots 4K120" is not, because the retrieved passage may not include the sentence that defined "it."
  2. State numbers out loud. A figure that only exists as an on-screen graphic does not exist to a language model. If it matters, the voiceover says it.

4. Description Length Was the One Metric That Correlated

Of every variable the study tested, description length was the only one with a meaningful positive relationship to citation frequency — a Pearson's r of 0.31. Weak-to-moderate, but it is the only signal in the dataset that pointed anywhere at all while views, likes, and subscribers sat at zero.

The plausible explanation: a long description is a second document about the video, in text, that the crawler can read without touching the transcript. It restates the claims, spells the proper nouns, and links out to sources.

So the description stops being an afterthought:

BlockWhat goes in itWhy
First 2–3 sentencesThe video's core answer, in plain proseThis is the passage most likely to be lifted
Chapter listTimestamps with query-shaped titlesFeeds deep-link citation
Summary section100–200 words restating the key claimsThe second readable document
SourcesReal outbound links to what you citedVerifiability, which retrieval systems reward
Links / CTAsEverything commercialLast, where it belongs

Keyword stuffing still does nothing. This is not a return to 2014 tag-spam — the classic ranking mechanics are in our YouTube SEO framework, and they have not been repealed. This is an addition, not a replacement.

5. If the Video Lives on Your Site, VideoObject Is Not Optional

Everything above assumes the video is on YouTube. For a video embedded on your own domain, Google has separate mechanics documented in Video SEO best practices and the VideoObject structured data reference. The requirements that trip people up:

  • The watch page has to be indexed and performing in Search before the video is eligible for video features at all. A great video on a page nobody can crawl is invisible.
  • The video must be embedded directly on the page, not hidden behind a tab, an accordion, or a click-to-load overlay.
  • It needs a thumbnail at a stable URL and a duration of at least 30 seconds.
  • JSON-LD is Google's preferred format, and Search Console's Video rich result report will tell you what is broken.

One page per video, with a real transcript in the HTML below the player. That transcript is the asset — the embed is just the delivery mechanism.

The Pre-Publish Gate

This runs on every long-form video before it goes out, and it takes about ten minutes:

  1. ✅ Title is phrased as the question a human would type or speak, not as a clever label.
  2. ✅ First 30 seconds of script state the answer in one sentence.
  3. ✅ Six to ten chapters, each titled as a query, each opening with its own conclusion.
  4. ✅ Human-corrected .srt uploaded — proper nouns, numbers, and product names verified letter by letter.
  5. ✅ Description: answer paragraph, chapter list, 100–200-word summary, real source links.
  6. ✅ Every number spoken aloud, not just shown on screen.
  7. ✅ Full nouns used before pronouns in any passage carrying a factual claim.
  8. ✅ Off-YouTube embeds: one page per video, VideoObject JSON-LD, transcript in the HTML.

Steps 1 through 3 are decided at the script stage, not in post. That is why this belongs in the brief — the structure that makes a video citable is the same structure that makes it retain, which is the argument in our scripting and hooks guide.

A Caveat Worth Stating Plainly

Being cited is not automatically good business. When AI Overviews began citing YouTube more often than medical websites for health queries — a finding drawn from a study of more than 50,000 health searches — it revealed a system that is confident about source selection in ways the underlying evidence does not always support. A citation is a machine's judgement that your passage answered a query. It is not a quality award, and it does not always convert to a view.

Track it anyway. Citation share is now a real distribution channel sitting alongside search, suggested, and browse, and it is not visible in any of the KPIs we usually watch — the ones that still matter are in our breakdown of YouTube analytics that count.

The Bottom Line

AI search engines cite structure, not popularity. Long-form video, chaptered into query-shaped sections, with a human-corrected transcript and a description that restates the claims in plain text, will get cited by a 400-view channel while a million-view video with auto-captions and a three-line description gets skipped.

None of this requires a different camera or a bigger budget. It requires deciding, at the script stage, that each section of your video has to make sense to someone — or something — that arrived in the middle.

👉 Start Your Project Now

Frequently asked questions

How do AI search engines like ChatGPT and Perplexity decide which videos to cite?
They read text, not video. An AI engine ingests the transcript, title, description, and metadata, then selects passages that answer the query directly. Popularity signals barely matter: a 2026 study of 100 million citation instances found the correlation between citation frequency and view count, likes, or subscriber count was effectively zero.
Do YouTube Shorts get cited by AI search?
Rarely. In Otterly's 2026 citation study, 94% of YouTube citations went to long-form videos and only 5.7% to Shorts. A short clip contains one claim with no surrounding structure, while a long-form video contains many self-contained passages a retrieval system can extract. Shorts remain useful for discovery, not for citation.
Do video chapters help with AI search visibility?
Yes, substantially. Among cited videos that carried timestamp signals, 78% were cited multiple times across two to five different chapters of the same video. Chapters turn one video into several independently retrievable documents. Title each chapter as a question someone would actually ask, and open each one by stating its answer.
Are auto-generated YouTube captions good enough for AI search?
No. Auto-captions mis-transcribe exactly the words that make a video worth citing: product names, technical terms, numbers, and proper nouns. Since the transcript is the only part of a video an AI engine can read, an error there becomes an error in the retrieved passage. Upload a human-corrected SRT file instead.
Does video description length affect AI citations?
It appears to. Description length was the only variable in the 2026 Otterly study with a meaningful positive relationship to citation frequency, at a Pearson's r of 0.31. A long description functions as a second readable document about the video, restating the claims in plain text with correct spellings and outbound source links.
What structured data does a video on my own website need?
VideoObject markup in JSON-LD, placed on the page where the video can be watched. Google also requires that the watch page itself is indexed and performing in Search, that the video is embedded directly rather than hidden behind a tab or overlay, that the thumbnail sits at a stable URL, and that the video runs at least 30 seconds.
Does getting cited by an AI search engine bring in views?
Not reliably. A citation means a model judged your passage the best available answer to a query, which is a distribution signal rather than a traffic guarantee, and many cited videos have very small audiences. Treat citation share as a channel worth tracking alongside search, suggested, and browse rather than as a replacement for them.