Long-Form Queries Rank 37 Positions Better on One B2B Security Site

The direct answer: across 90 days of search data reaching one B2B cybersecurity site, long and conversational queries ranked at an impression-weighted position of 10.0, while short keyword head terms averaged 46.9. Median positions were 8.0 and 43.0. That gap survives every robustness check we ran. Neither group produced meaningful clicks, and the difference in clicks between them is not statistically distinguishable, which is a separate and more awkward finding.
Data pulled August 25, 2026, analyzed and corrected August 28, 2026. Method, raw counts, robustness checks and limitations are stated in full below.
Most writing about AI search is speculation about what engines might do. This is a small, checkable measurement of what actually reached one site. We publish the method alongside the numbers so anyone can run the same classification against their own Search Console property.
Method
The source is Google Search Console for a single domain-level property, ztekcyber.com, covering 2026-05-28 through 2026-08-25. We pulled the complete query export with pagination, 1,339 rows, then removed 8 rows containing search operators such as -site: (scraper artifacts rather than queries) and 10 brand queries containing the company name or its acronyms. That leaves 1,321 queries and 15,204 impressions. Those 1,321 queries are every non-brand query Search Console will disclose for the window.
Every query was assigned to one class by text shape alone:
| Class | Rule | Example from the data |
|---|---|---|
| Long-form | Opens with a question word, or contains a first-person marker, or runs to 8+ words | "what are the best vciso services for mid-market organizations?" |
| Head term | Four words or fewer, no question word, no first-person marker | "virtual ciso services" |
| Mid-tail | Everything between those two definitions | "cybersecurity metrics for the board" |
Position figures are impression-weighted, so a query with 300 impressions counts for more than one with three. Medians are reported alongside, because weighting makes a single large query influential and the reader should be able to see both.
Result 1: the position gap is 37 places
| Class | Queries | Impressions | Share | Weighted pos. | Median pos. | In top 10 | Clicks |
|---|---|---|---|---|---|---|---|
| Long-form | 287 | 3,521 | 23.2% | 10.0 | 8.0 | 88.1% | 0 |
| Head term | 620 | 8,713 | 57.3% | 46.9 | 43.0 | 6.4% | 7 |
The specific head terms make the pattern concrete. "virtual ciso consulting services" drew 595 impressions at position 62.1. "virtual ciso services" drew 349 at 62.9. "nist csf assessment" drew 268 at 54.1. Those are the terms a security firm would name if you asked it to list its keywords, and they are all effectively unreachable. Meanwhile a seventeen-word question about whether a named competitor supports multi-framework compliance sat at position 5.5.
Robustness
One query, a twelve-word NERC CIP requirement string, supplies 2,053 impressions, which is 58.3 percent of the long-form group. That concentration is the biggest threat to the headline number, so here is the same comparison under three cuts:
| Cut | Long-form weighted pos. | Gap vs head terms |
|---|---|---|
| All long-form queries | 10.0 | 36.9 |
| Excluding the dominant query | 13.1 | 33.8 |
| Question and first-person only, no length rule | 14.0 | 32.9 |
The gap narrows but does not disappear, and it stays above 30 positions under every cut. That is why we are willing to report it at all.
Result 2: most of the group is long, not conversational
An honest description of the classifier matters more than a flattering one. Of the 3,521 impressions in the long-form group, 77 percent qualify on the eight-word rule alone, with no question word and no first-person marker. The largest member is a NERC CIP requirement string. The second largest, "ai supply chain risk monitoring cybersecurity third party", is a bare keyword pile.
So this is not cleanly a study of natural-language questions. It is a study of long queries, inside which sits a smaller genuinely conversational set of 160 queries and 813 impressions. That subgroup ranks at weighted position 14.0 and median 8.0, still roughly 33 positions ahead of head terms. Both readings support the same conclusion, but "long" is the honest label for the main group.
Result 3: the share depends entirely on where you draw the line
Any claim of the form "X percent of search is now AI-shaped" is a claim about a threshold, so here is the full sensitivity rather than the single number that best supports an argument:
| Definition | Share of non-brand impressions | Weighted position |
|---|---|---|
| Question word or first person only | 5.3% | 14.0 |
| Plus 8 or more words (used above) | 23.2% | 10.0 |
| Plus 6 or more words | 32.7% | 14.4 |
| Plus 5 or more words | 42.7% | 18.4 |
Between roughly 5 and 43 percent of non-brand impressions look long or conversational depending on strictness, and the ranking advantage holds at every threshold.
Want this run against your own property?
A Z Cyber advisor can walk your Search Console data through the same classification and show you which shapes you already win.
Result 4: some queries arrive with their template variables unfilled
Three queries in the dataset contain literal placeholder tokens no person would type. The largest, at 31 impressions and position 12.0, reproduced exactly as it appears in the export:
as a vp it risk and compliance at a enterprise_1000_plus company in united_states,compare secureframe, sprinto, scytale, vanta, drata, thoropass, and scrut automation on scalability for managing multi-framework compliance. you must provide a forced ranking from best to worst.
Two more follow the same construction with mid_market_500_1000, one from the Netherlands and one from the United States. The underscore-delimited tokens are variable names from a prompt template that was never populated before the request went out. The missing space after the comma is part of the same artifact: the template concatenated its segments without one. Both quotes on this page are reproduced verbatim, including that.
The instruction at the end is the giveaway about intent. It does not ask for an assessment, it demands a forced ranking from best to worst across seven named vendors, which is a benchmarking pattern rather than a buying question. Three queries and 33 impressions is a tiny sample and we are not building a trend on it. It is worth reporting because it is direct evidence, visible in an ordinary Search Console export, that machine-generated comparison prompts reach vendor sites.
Result 5: almost nothing clicks, and the two groups cannot be told apart
This is where an earlier version of this post overreached, so it is worth being precise. Long-form queries produced zero clicks on 3,521 impressions. Head terms produced seven on 8,713, a rate of 0.08 percent. At that rate, the expected number of clicks on 3,521 impressions is 2.8, and observing zero is unremarkable: a Fisher exact test returns p equal to 0.20. The two groups are not statistically distinguishable on clicks, and any claim that one converts worse than the other is unsupported by this data.
What the data does support is duller and still useful. Ranking well on this site produced almost no clicks for anybody. A page can sit at position 8 for hundreds of impressions and register nothing, which is the expected pattern when answers are synthesized on the results page, but Search Console cannot see inside a results page or an assistant, so this data cannot prove that a citation occurred or that a reader was influenced.
The practical consequence is a measurement gap rather than a marketing claim. If you publish for this shape of query, clickthrough rate will not tell you whether it worked, and you need a second instrument that samples assistant answers directly and records whether you were named.
Limitations, stated plainly
Concentration. One query is 58.3 percent of the long-form group's impressions. Every headline figure is reported above alongside cuts that remove it.
Coverage. Query-level data accounts for 15,689 of the site's 46,143 total impressions in the window. Google anonymizes rare queries, and rare is what long conversational queries tend to be, so the long-form share here is more likely understated than overstated.
Confounding. Three explanations fit the position gap and we cannot separate them. Competition is the simplest: a four-word commercial head term has far more established pages contesting it than a seventeen-word question does. Content fit is at least as plausible: this site publishes long informational and vendor-comparison articles and does not have strong commercial service pages for "ciso consulting", so the comparison is partly between queries the existing content answers and queries it does not. The two groups are also not the same subject matter, since the long-form set skews toward GRC platform comparison while the head-term set skews toward vCISO consulting. Engine preference is the third possibility and the one we consider least likely.
Attribution. Query shape is inferred from text. Apart from the three template-variable cases, no query can be attributed with confidence to a specific assistant or to automation. A person typing a full question and an assistant issuing one look identical here.
Scope. One site, one vertical, one 90-day window. This is a case study. It should not be cited as a market statistic, and the numbers above describe ztekcyber.com rather than B2B search in general.
How to run this on your own site
Export your full query list from Search Console for the last 90 days, using pagination rather than a capped row limit. A truncated export can silently sort alphabetically and censor your sample, which is exactly the error the first version of this analysis made. Drop brand queries and anything containing a search operator. Tag each remaining query as long-form if it starts with a question word, contains a first-person marker, or runs to eight or more words. Compute impression-weighted and median position for that group and for queries of four words or fewer, then check whether your result survives removing your single largest query.
If the pattern holds on your property, the implication is the same one it had for us. The terms you would have picked as keywords are probably unreachable, the shapes you rank for are questions you may not have deliberately targeted, and your clickthrough column will not tell you either way.
What to watch next
We intend to re-run this each quarter against the same property and publish whether the share and the gap move. The next run covers the window ending in late November 2026. Two things would change our reading: a sustained narrowing of the position gap as more publishers write for long-form shapes, or a first non-zero click volume on long-form queries, which would suggest engines are sending traffic rather than absorbing it. Until then, treat the numbers above as one site's honest snapshot. If you want to compare notes on what your own property shows, talk to a Z Cyber advisor.
Frequently Asked Questions
Do long-form search queries rank better than short keyword terms?
On the site we measured, substantially better. Across 90 days and 1,321 non-brand queries, queries that were long or conversational carried an impression-weighted position of 10.0, while head terms of four words or fewer averaged 46.9. Median positions were 8.0 and 43.0. This is one site in one vertical, so treat it as a case study rather than a market finding.
What counts as a long-form or AI-shaped query?
In this study a query was classified as long-form if it opened with a question word, contained a first-person marker such as my or we, or ran to eight words or more. That last rule does most of the work: 77 percent of the impressions in the group qualified on length alone and are keyword strings rather than sentences. Short strings like virtual ciso services are head terms.
Do high-ranking pages for conversational queries get clicks?
In this dataset almost nothing clicked either way. Long-form queries produced zero clicks on 3,521 impressions and head terms produced seven on 8,713, a rate of 0.08 percent. A Fisher exact test on that comparison returns p equal to 0.20, so the two groups cannot be distinguished on clicks. The useful observation is that ranking well produced almost no clicks for either group, not that one outperformed the other.
How can you tell that some search queries are generated by AI tools?
Some arrive with their template variables still unfilled. Three queries in this dataset contained literal placeholder tokens such as enterprise_1000_plus, mid_market_500_1000 and united_states, inside sentences that then demand a forced ranking of named compliance vendors. A person does not type an underscore-delimited variable name. These are machine-generated prompts from an automated harness, reaching search logs verbatim.
What are the limitations of this search query study?
It covers one site in one vertical over one 90-day window, so it is a case study and not a market survey. Query-level data accounts for 15,689 of the site's 46,143 impressions because Google anonymizes rare queries. A single query supplies 58 percent of the long-form group's impressions, so headline figures are reported alongside robustness checks that exclude it. Query shape is inferred from text, so attribution to any specific assistant is not possible.
Subscribe for Updates
Get cybersecurity insights delivered to your inbox.


