Introduction: The Structure Underneath the Words
Semantic geometry is not a metaphor. It is the measurable spatial relationship between your content’s HTML structure, information density, and meaning boundaries that determines whether AI systems can parse, chunk, and cite what you publish. When the geometry is wrong, AI cannot isolate a discrete knowledge unit from your page, and no citation follows. When the geometry is right, your content becomes a pre-validated inference node that LLMs retrieve with high confidence. This structural logic extends beyond individual pages—horizontal content analysis reveals how the same geometric principles expose bias and structural failure across comparison-format content.
This post is the canonical definition of that term. Everything else on this site, from schema to linking to symmetry, operates on top of it.
Why Ranking at Position 1 Still Earns Zero Clicks
We watched this happen in our own Search Console data: a page can hold position 1 in Google’s classic index for its defining query and still receive zero clicks, because an AI Overview intercepts the query and assembles its answer from other sources. The ranking system and the citation system are no longer the same system. One rewards a page. The other retrieves a knowledge unit: a discrete, self-contained container it can lift and quote. Until now, “semantic geometry” lived inside our technical hub as a supporting concept, never as a dedicated definitional container. With no discrete unit to retrieve, the AI defaulted to the statistical consensus around the term and answered without us. That is exactly the Average Answer mechanism we have documented elsewhere. This post exists to close that gap: it is the origin node for the term.
In the first eight-weeks of our twelve-week AI Citation Study, we see this split every week. On one of our client sites, the page that answers “how MDR strengthens cyber risk management” sits at position 3 in Search Console for that query with zero clicks — and is the source Google’s AI answer and Perplexity both cite when a buyer asks the question. Ranking put the page in the room; structure is what got it quoted. Across seven of the eight client sites we track, pages Google was already testing on page one — even with no clicks at all — went on to be cited at roughly 1.5 to 2.6 times the rate of pages it wasn’t testing. Being in the pool matters. It is not the same as being cited, and the click count tells you nothing about which one you are.
Semantic Geometry, Defined in One Paragraph
VizzEx Semantic Geometry is the spatial and structural relationship between three things: your HTML hierarchy (how header levels nest), your content density (how much information sits inside each structural container), and your meaning boundaries (where the machine can tell that one knowledge unit ends and another begins). AI retrieval systems evaluate these three properties before they evaluate meaning. Geometry is the pre-meaning layer: the structural precondition that must be satisfied before any semantic relationship in your content can be recognized, weighted, or cited.
What Is Semantic Geometry?
The Definition, Expanded
In the VizzEx Signal Architecture framework, semantic geometry describes content as a physical structure with measurable properties, not as prose with decoration. Put plainly, it’s content structure for AI: where traditional on-page SEO optimizes words for a ranking algorithm, semantic geometry optimizes structure for the chunking and embedding pipeline that runs before meaning is evaluated at all. A page is a set of containers. Headers declare what each container holds. Density determines whether the container’s contents can be summarized into a single clean mathematical representation. Boundaries determine whether the container can be lifted out of the page intact. Every downstream AI process (embedding, retrieval, synthesis, citation) inherits the quality of those three properties.
Why “Geometry” Is the Right Metaphor
The word matters. Meaning in an LLM is literally geometric: content is converted into vector embeddings (coordinates in high-dimensional space), and retrieval is a distance calculation between a query’s coordinates and your content’s coordinates. A well-structured section produces a sharp vector: a discrete shape with clear edges that lands precisely in semantic space. An unstructured 1,500-word wall of text produces an amorphous blob: a single muddy vector trying to average fifteen distinct ideas into one location. Discrete shapes get retrieved. Blobs get skipped.
How AI Systems Read Geometry at Inference Time
Before a RAG (Retrieval-Augmented Generation) pipeline ever runs AI semantic analysis, the machine evaluation of what your page means, a chunker mechanically slices the page into segments, commonly 256 to 512 tokens, roughly 150 to 300 words. That chunker doesn’t read for comprehension. It reads structure: header positions, sectioning elements, DOM boundaries. This is semantic chunking in its rawest form, and it happens to every page, on every fetch. Your HTML is the cutting guide. If the guide is clear, every chunk is a coherent knowledge unit. If the guide is absent, the machine cuts mid-thought, and every resulting fragment is damaged inventory.
The Three Axes of Semantic Geometry
Axis 1: Vertical Hierarchy (Headers as Meaning Containers)
The first axis is vertical: the H1 → H2 → H3 cascade. To an AI parser, header levels aren’t typography. They’re a containment declaration. An H2 states: “everything below me, until the next H2, belongs to this concept.” An H3 nests a sub-concept inside it. This is how the machine reconstructs your logic tree without reading a single sentence. A page whose headers skip levels, repeat keywords without differentiation, or style paragraphs as fake headings hands the parser a corrupted tree, and a corrupted tree can’t be traversed with confidence.
Axis 2: Horizontal Density (The 1:100 Ratio)
The second axis is horizontal: how much information sits under each header. The 1:100 baseline comes from Carolyn Holzman’s multi-year forensic SEO research: 20 to 25 test pages published every month across 4.5 years of daily indexation observation, work she has since extended and sharpened as a Partner at VizzEx. That testing surfaced an empirical Goldilocks Zone: content with roughly one header per 100 words (a ~10% header-to-word ratio) indexed consistently, while content below that ratio was hit-or-miss and content pushed past ~18% also failed. The full methodology is documented in how the 1:100 header-to-word ratio was discovered and validated. The ratio isn’t a style preference. It’s preemptive chunking: a header every ~100 words hands the RAG pipeline segments already sized to its own default chunk window.
Axis 3: Boundary Confidence (Where Knowledge Units End)
The third axis is boundary confidence: how certainly the machine can determine where one knowledge unit stops and the next begins. HTML5 sectioning elements do this work explicitly. A <section> mathematically defines a header’s Scope of Relevance: where its influence ends. An <article> declares “everything inside here is the primary unique knowledge.” An <aside> declares “related, but weight it less.” Without these signals, the algorithm must guess at boundaries, and every guess lowers the confidence score attached to any chunk it extracts.
Why Semantic Geometry Determines AI Citability
The RAG Pre-Processing Step Most SEOs Ignore
Here’s the step most SEOs never consider: the AI slices your page apart before it ever decides whether you’re relevant. Legacy SEO was built for a system that read whole pages and ranked whole URLs. AI retrieval does neither. It chops your page into chunks, turns each chunk into a vector, and retrieves and cites chunks. Not pages. Chunks. So if your structure fails at the chopping stage, you’re out before relevance is ever calculated. You can have the best answer on the internet and still lose, because the machine never carried you into the room where the comparison happens.
Where the cited passages actually come from
If the machine retrieved pages, where a passage sits wouldn’t matter. It retrieves chunks, and you can see it in what Google shows. Google’s AI answers return the exact passage they lifted from a cited page; we located 928 of those passages on pages in our study and recorded where in the page each one came from.
| Position in the page | Sites built to the geometry (891 passages, 8 sites) | Two famous marketing publishers (37 passages) |
|---|---|---|
| First 10% of the page | 25% | 65% |
| First 20% | 38% | 76% |
| Middle 30–80% | 36% | 14% |
| Last 20% | 15% | 8% |
| From the Q&A block at the foot of the post (pages that have one) | 20% of passages, from 24% of the text | 1 passage in 31 |
On the two famous publishers’ pages, three-quarters of the passages came from the top fifth of the page — the “key takeaways” box, the numbered list at the head of a listicle. The rest of the page was, for citation purposes, nearly invisible. On sites built to the geometry, passages came from everywhere: 38 percent from the top fifth, 36 percent from the middle of the page, 15 percent from the last fifth — and the question-and-answer block that sits at the very foot of every treated post supplied passages in proportion to its length, one in five. Every section that passes the standalone-quote test is a candidate. Sections that don’t, aren’t, no matter how high on the page they sit.
Noisy Data In, Low-Confidence Embeddings Out
AI pipelines are data-quality engines before they are answer engines. A section that wanders across four ideas without a boundary produces an embedding that represents none of them well. This is what the geometry framework calls a muddy vector. Muddy vectors sit far from any specific query’s coordinates, so they lose every retrieval distance calculation to sharper competitors. The fix isn’t better writing within the blob. The fix is geometry: breaking the blob into discrete, single-concept containers so each embedding lands exactly where its meaning lives. Vector embedding quality is downstream of container quality, always.
Boundary Confidence and Cross-Entropy Validation
Boundary quality is measurable. Cross-Entropy Validation (CEV) is the instrument VizzEx uses to score whether a section’s declared structure and its actual content agree. In effect, it measures how “surprised” a model is by what it finds inside a container relative to what the container’s header and schema promised. Low surprise means high boundary confidence and a chunk the model can quote standalone. High surprise means the section fails adversarial verification. The mechanics are detailed in how cross-entropy validation scores geometric compliance at the model level.
The Empirical Evidence: Geometric Compliance Achieves 100% Indexation
The strongest evidence that geometry is causal, not cosmetic, came from the indexation testing program itself: once test pages were brought into geometric compliance (the 1:100 density ratio, clean vertical hierarchy, explicit sectioning), indexation reached 100%, against the hit-or-miss baseline of non-compliant pages published under identical domain conditions. Indexation is the first gate of the AI supply chain, and geometry moved that gate from probabilistic to deterministic.
Probationary Citations and the Flickering Problem
Clearing a gate once isn’t the same as settling behind it. Noisy pages do get pulled into citations, especially when retrieval is sampling widely for a query with no settled origin node. Those are probationary citations: the model cites the page before it has thoroughly examined it. Once deeper verification runs (symmetry checks, boundary scoring, re-chunking on the next fetch), the noisy source rotates out and another rotates in. We call this flickering: citations that appear, vanish, and reappear somewhere else because the AI is unsettled about the stability of what it’s citing. Flickering isn’t visibility. It’s the audit happening in public. Geometry doesn’t just get you through the first gate; it’s what makes a citation stick.
Probation, measured on four engines
Probation is measurable. In our study, a page cited for the first time on a question is still there at the next weekly check about half the time — 48 percent on Google’s AI answers, 49 percent on Claude, 49 percent on Perplexity — and just 17 percent on ChatGPT. A page that was already being cited the week before survives at 71, 85, 83 and 50 percent.
| Engine | First citation on a question — still cited next check | Already-cited page — still cited next check |
|---|---|---|
| Google AI answers | 48% (n=5,362) | 71% (n=6,515) |
| Claude | 49% (n=1,549) | 85% (n=4,987) |
| Perplexity | 49% (n=3,419) | 83% (n=13,395) |
| ChatGPT | 17% (n=5,605) | 50% (n=1,999) |
Two checks out, half of first citations are gone on every engine. That is the flicker: the model tries a source, looks again, and more often than not puts it back. What the numbers also say is that nothing you can see in the page’s markup predicts which new citations survive — pages that pass a structural audit with a perfect score are confirmed at the same rate as pages that fail it. Confirmation is decided somewhere the markup isn’t. That deciding factor is the question itself—understanding whether AI cites or recommends your content depends on query intent in ways that geometry alone cannot control.
Semantic Geometry vs. Traditional SEO Structure
Keyword Density vs. Information Density
Traditional SEO measured keyword density: how often a target phrase appears per hundred words. Semantic geometry measures information density: how many discrete, self-contained ideas sit inside each structural container. These are orthogonal, and optimizing the first frequently destroys the second. A page can hit every keyword target while packing five unrelated claims under one header, producing a single unretrievable blob. The machine isn’t counting your phrases. It’s testing whether each container holds exactly one clean, quotable unit of knowledge. And a clean container is only half the test: the unit also has to carry its explicit semantic relationships, the links that declare how this idea connects to your other ideas and to verified outside sources. A chunk with no declared relationships is an orphan the AI can extract but has no reason to trust. Geometry makes the unit liftable; relationships make it citable.
Why Header Keyword Stuffing Destroys Geometric Integrity
The legacy reflex of stuffing target keywords into every H2 actively corrupts the vertical axis. When six headers on a page all approximate the same phrase, the containment declaration collapses: the parser can no longer distinguish which container holds which concept, and the logic tree flattens into repetition. Headers exist to differentiate containers, not to repeat the page’s topic six times. A geometrically compliant header answers one question: “what does this container hold that no other container on this page holds?”
The Flat-List Penalty: Why Bullet-Heavy Pages Fail AI Retrieval
The inverse failure is the bullet-heavy, heading-light page: long runs of list items with no containing structure. Lists feel organized to human eyes, but to a chunker a hundred bullets under one header is still one container: one muddy vector with a hundred fragments inside it and no boundaries between concepts. Bullets are furniture inside a room; they’re not walls. Geometry requires walls: headers and sectioning elements that give each cluster of related items its own addressable, extractable container.
Semantic Geometry as a Signal Architecture Component
Where Geometry Sits in GEOMesh
In the VizzEx GEOMesh (Generative Engine Ecosystem) architecture, semantic geometry is the foundation layer. GEOMesh maps a domain as a connected knowledge graph (categories, hubs, semantic relationship links, relational schema), but every one of those higher-order structures assumes the page-level containers beneath them are sound. A semantic relationship link pointing into a geometrically malformed page delivers the AI to a blob it can’t parse. Site-level architecture can’t compensate for page-level geometry, any more than a road network compensates for collapsed buildings.
The Symmetry Gate: Geometry’s Pass/Fail Checkpoint
Under the VizzEx Signal Architecture Protocol v1.2, the Symmetry Gate™ is the binary validation that your geometric claims are true: 1:1 parity between the raw HTML source, the rendered DOM, and the declared schema. A page whose headers exist only after client-side JavaScript executes, or whose schema declares structure the DOM doesn’t physically contain, throws a Symmetry Failure and falls through the gate before geometry is even scored. Compliance is binary and sequenced: first the structure must be honest, then it can be measured.
What eight weeks of head-to-heads show
We can put a number on the gate. In the first eight-weeks of our twelve week AI Citation Study, every tracked page is scored weekly against the pages currently winning the AI answer on the same question — same audit, same fifty measures, both sides — and the win rate steps with how much of the geometry is in place.
| What the site carries | Structural head-to-head win rate vs. the pages currently holding the AI answer |
|---|---|
| Full methodology — geometry, relational schema, semantic-relationship links (three client sites, vizzex.ai, and my company’s site) | 85–94% |
| Semantic-relationship links only, no declaration layer (one client site) | 76% |
| Mid-rollout — geometry on roughly a third of the library at the time of measurement (one client site) | 55% |
| Full stack on templates that break raw-to-rendered symmetry (one client site) | 49% |
| Two famous marketing publishers, professionally run, no geometry work | 12–14% |
Sites built to the full geometry win those head-to-heads 85 to 94 percent of the time. Two famous marketing publishers, with every authority signal the industry says matters, win 12 to 14 percent. Between them sit the partial cases, and they sit exactly where the geometry predicts: a site with the semantic link layer but no declaration at 76 percent; a site midway through its rollout at 55; and a site with the full stack installed on templates that break symmetry between source and rendered DOM at 49 — the declaration is there, the page can’t confirm it, and the gate does what it is built to do. Each of those structural layers can be evaluated independently using Chunk Autonomy applied as an editorial test — the same principle that explains why a passage must stand alone before the architecture around it can matter.
The staircase is not a comparison of different companies. When the mid-rollout site’s next batch of pages was armored in September, its win rate against the answer-holders moved from 58 to 68 percent at the following weekly check — same site, same questions, more of the library compliant. The compounding effect is even more pronounced when the agentic loop passes through a retrieval layer rather than re-reasoning from scratch — an architectural distinction that determines how much of your structural investment survives the compute pipeline.
Geometry Enables Attribution Weight
Semantic relationship links, the explicit semantic connections that demonstrate how your ideas relate, only carry full attribution weight between geometrically sound endpoints. When both the linking chunk and the destination chunk are discrete, boundary-confident knowledge units, the AI can ingest the relationship as a clean node-to-node record. This is also why flat schema actively undermines geometric signal integrity: isolated metadata blocks with no physical anchor into the page’s geometric containers declare relationships the machine can’t verify, which reads as risk, not richness.
Geometry First, Schema Second
A common confusion is whether semantic geometry is just another name for structured data. It isn’t, and the layers run in a fixed order. Schema markup declares what your content is; semantic geometry determines whether the machine can verify that declaration against the physical page. Schema attached to a geometrically malformed page fails validation: the claims can’t be anchored to discrete DOM containers, which triggers asymmetry flags rather than trust. Geometry is the structural precondition. Relational schema is the declaration layer built on top of it, and it only carries weight when the containers underneath it are real.
Geometry Must Be Horizontal, Not Heroic
One perfectly structured post on a geometrically chaotic site is a statistical outlier, not a signal. AI systems evaluate domain-level topical coherence, which means geometric compliance has to hold horizontally, across the full content library, before the domain reads as a trustworthy knowledge system. This is why horizontal content analysis, not post-by-post editing, is the diagnostic that matches how the machine actually evaluates you. This is the structural layer of how semantic content analysis maps relationships across a geometrically structured site: the relational map is only as reliable as the geometry of the pages it connects.
The one place the data surprised us
“Horizontal, not heroic” is the one place our data surprised us. On the sites built to the geometry across the whole library, structure stopped predicting which of their pages got cited — because pages that have never been cited on those sites still win roughly 90 percent of their structural head-to-heads. Structure had become the floor, not the differentiator; what decided the citation from there was the question, the entity, and the links. That retention, however, is not permanent — understanding citation half-life and signal eviction mechanics explains why even well-structured pages must be maintained to hold their position in the answer layer.
The exception proves the rule. On the one client site midway through its rollout, only a third of pages sit in the density band, the site as a whole wins 55 percent of head-to-heads against the answer-holders instead of 85 to 94, and its un-armored pages are the ones still losing. A perfectly structured post on a half-structured site is an outlier the model has no reason to trust.
How to Diagnose and Fix Semantic Geometry Problems
Step 1: Audit Header-to-Word Ratios Per Section
Start with the horizontal axis, and measure at the right resolution: per section, not per page. A page-level average of 1:100 can hide a 700-word blob in the middle of an otherwise compliant post, and the blob is precisely the section that will fail retrieval. Walk each container: count the words between each header and the next. Anything drifting far past ~150 words is a candidate for subdivision; anything under ~40 words repeatedly may be pushing toward the over-fragmentation ceiling near 18%.
Step 2: Identify Orphaned Sections
Next, audit the vertical axis. An orphaned section is content with no clear parent concept container: an H3 whose parent H2 doesn’t logically contain it, or a stretch of prose sitting between the end of one section’s idea and the start of the next header’s territory. Orphans are boundary ambiguity in its purest form: the parser can’t decide which container owns the content, so the content dilutes both adjacent vectors. Every paragraph on the page should have exactly one unambiguous structural parent.
Step 3: The Standalone Quote Test
The fastest boundary-confidence diagnostic requires no tooling: read any single section in total isolation and ask, could an AI quote this standalone? Does it name its subject explicitly rather than leaning on “this” and “it” resolved three sections earlier? Does it carry its own claim, its own evidence, its own attribution? This is Chunk Autonomy applied as an editorial test. If a torn-out section can’t survive alone, it won’t be cited alone. And chunks are only ever cited alone.
Step 4: Align DOM Output With Rendered HTML
Finally, verify that your geometry physically exists where the machine looks for it. Headers, sectioning elements, and schema anchors must be present in the server-rendered HTML, not assembled client-side by JavaScript after load. A crawler parsing raw source must see the identical structure a rendering agent sees, or the mismatch triggers the Symmetry Failure described above. Server-rendering isn’t a performance nicety here; it is the delivery mechanism that makes your geometric claims verifiable at zero compute cost.
How VizzEx Pro Automates Semantic Geometry Audits at Scale
The four steps above can be run by hand on a single post. Across a full library, VizzEx Pro runs them horizontally: per-section header-to-word ratios against the 1:100 baseline, vertical hierarchy integrity, HTML/DOM symmetry under the VizzEx Symmetry Gate™ standard, and boundary confidence scored via the Cross-Entropy Validation (CEV) baseline. Results roll up into the GEOMesh architecture view, which shows which pages function as clean inference nodes and which sections fail chunking before their content is ever evaluated.
Semantic Geometry and the AI Citation Half-Life
Compliant Geometry Refreshes Faster
Citations decay. But whether they disappear is a matter of architecture. The observed citation half-life runs roughly 4.5 weeks, a finding from outside research we break down in our citation decay analysis and are now testing against our own tracking: an AI model’s confidence in a cached source degrades on a metabolic cycle, and the source must be re-verified (re-fetched, re-chunked, re-validated) to hold its position in the answer layer. Geometrically compliant content is computationally cheap to re-verify: clean containers, honest symmetry, zero-surprise boundaries. Cheap verification means frequent re-induction. Expensive verification means the model quietly substitutes a cheaper source, and your citation slot is gone.
The Compute Tax Is the Enforcement Mechanism
The economics enforcing all of this are blunt. Rendering a JavaScript-dependent, geometrically ambiguous page can cost up to 1,000x more CPU cycles than parsing clean static HTML. That is the VizzEx Compute Tax. An AI agent facing a malformed page must also fire repeated agentic loop passes (plan, fetch, criticize, re-fetch) to resolve what a compliant page states outright. Retrieval systems are cost-benefit engines: they systematically route around expensive sources. Good geometry is, in the end, a Compute Bribe: you make citing you the cheapest available option.
The enforcement is visible when it fails
Early in our study, one client site with the full geometry installed was, unknown to anyone, returning 403s to the AI crawlers — the pages were perfect and the machines couldn’t fetch them. Of 29 tracked positions, zero held until the block was found and lifted. In the same window, a second client site with the full stack installed on templates that break raw-to-rendered symmetry held zero of 25.
Maximum structure, impaired delivery, no retention. The geometry has to be served — in the source HTML, to the bot, at the cost of a plain parse — or it is not geometry the machine ever sees. The same failure mode appears when headers exist only after client-side JavaScript renders them — the structure is real to a browser and invisible to every AI crawler that never executes the script.
Conclusion: Semantic Geometry Is Not Optional for AI Citability
The Three Axes, Restated
Vertical hierarchy tells the machine what contains what. Horizontal density, the 1:100 ratio, keeps every container within the chunk window that retrieval systems actually process. Boundary confidence lets each container be lifted out, quoted, and attributed standalone. Fail any axis and the failure cascades: bad geometry corrupts your semantic signals at the source, corrupted signals produce muddy embeddings, muddy embeddings lose retrieval, and lost retrieval means no citation, regardless of how authoritative the words inside the structure were. Semantic clarity begins as a structural property before it is ever a writing property.
Find Your Geometric Failures Before the Models Do
Every AI system that touches your content runs this structural evaluation silently, on every fetch, with no error report sent back to you. The only way to see your geometry the way the machine sees it is to audit horizontally (every post, every section, every boundary) before the next induction cycle prices your domain. Run a horizontal blog analysis with VizzEx Pro and find the sections that are failing chunking today, while they’re still yours to fix.
Frequently Asked Questions
What is semantic geometry and how does it affect AI citability?
VizzEx Semantic Geometry is the spatial and structural relationship between three things: your HTML hierarchy (how header levels nest), your content density (how much information sits inside each structural container), and your meaning boundaries (where the machine can tell that one knowledge unit ends and another begins). AI retrieval systems evaluate these three properties before they evaluate meaning. Geometry is the pre-meaning layer: the structural precondition that must be satisfied before any semantic relationship in your content can be recognized, weighted, or cited.
Why can a page rank #1 in Google and still get zero clicks or AI citations?
A page can hold position 1 in Google's classic index for its defining query and still receive zero clicks, because an AI Overview intercepts the query and assembles its answer from other sources. The ranking system and the citation system are no longer the same system. One rewards a page. The other retrieves a knowledge unit: a discrete, self-contained container it can lift and quote.
What is the 1:100 header-to-word ratio and why does it matter for AI retrieval?
The 1:100 baseline comes from Carolyn Holzman's multi-year forensic SEO research: 20 to 25 test pages published every month across 4.5 years of daily indexation observation. That testing surfaced an empirical Goldilocks Zone: content with roughly one header per 100 words (a ~10% header-to-word ratio) indexed consistently, while content below that ratio was hit-or-miss and content pushed past ~18% also failed. The ratio isn't a style preference. It's preemptive chunking: a header every ~100 words hands the RAG pipeline segments already sized to its own default chunk window.
How do AI systems chunk and process web page content before deciding what to cite?
Before a RAG (Retrieval-Augmented Generation) pipeline ever runs AI semantic analysis, a chunker mechanically slices the page into segments, commonly 256 to 512 tokens, roughly 150 to 300 words. That chunker doesn't read for comprehension. It reads structure: header positions, sectioning elements, DOM boundaries. Your HTML is the cutting guide. If the guide is clear, every chunk is a coherent knowledge unit. If the guide is absent, the machine cuts mid-thought, and every resulting fragment is damaged inventory.
What HTML elements should I use to define clear content boundaries for AI crawlers?
HTML5 sectioning elements do this work explicitly. A <section> mathematically defines a header's Scope of Relevance: where its influence ends. An <article> declares 'everything inside here is the primary unique knowledge.' An <aside> declares 'related, but weight it less.' Without these signals, the algorithm must guess at boundaries, and every guess lowers the confidence score attached to any chunk it extracts.