Search Engineering

Technical work for Discover, and what is not within your control

Discover has documented eligibility requirements and an unpublished ranking system. The technical work is real and it is bounded — and the part that determines traffic is not the part anyone can change.

The problem

Discover traffic is volatile for reasons outside the site

A publisher sees a month of Discover traffic and then a month without it, and the natural response is to look for the technical fault. Sometimes there is one — a missing image, a page that fails an eligibility requirement — and often there is not. Discover surfaces content it judges interesting to a particular user, from a feed that changes constantly, and no technical configuration makes a page appear in it. The work is to establish eligibility, remove anything that would exclude a page, and measure honestly.

  • Discover traffic appeared and then stopped, and the cause is not established
  • Articles are missing images in Discover, or the wrong image is used
  • Nobody can say whether the site meets Discover's documented requirements
  • The site is treating Discover as a channel it can optimise for directly
  • Discover traffic is being attributed to search in reporting
  • Large images are not being served at the sizes Discover uses
  • Traffic from Discover varies enormously month to month and is treated as a regression
  • The site is news-oriented and Discover is a significant share of its traffic
Who this is for

The people who usually bring us this problem

A publisher with Discover traffic

It is a real share of the total and its variability needs understanding rather than chasing.

A publisher whose articles look wrong in Discover

Images are missing, cropped badly or not the ones intended.

A publisher establishing whether it is eligible

The documented requirements need checking against the site rather than assumed.

What it costs

What this costs while it goes unfixed

Engineering faults are rarely confined to the engineering layer. These are the commercial consequences we see most often.

Inclusion is not something a site can configure

Discover selects content it judges likely to interest a user, from a system whose criteria are not published. The documented requirements are prerequisites — a page that fails them cannot be eligible — and meeting them does not produce inclusion. This distinction is the difference between work that is bounded and work that is not, and it is why no provider can offer Discover traffic.

Images are the technical prerequisite most often failed

Discover uses a large image, and the requirements around size, aspect ratio and availability are documented. A page whose image is too small, is loaded by JavaScript, or is blocked by a robots rule cannot be eligible — and the failure is not reported anywhere, so it presents as a page that simply does not appear.

Traffic is inherently variable and treating it as a regression is a mistake

A feed that changes per user and per session produces traffic that moves for reasons unrelated to the site. Month-on-month comparison of Discover is close to meaningless, and the useful measurements are eligibility, image handling and whether the pages that did surface were the intended ones.

Discover traffic is reported with search and is a different thing

It arrives in the same reporting and it is not search — it is a feed, driven by interest rather than by query. Treating it as part of search distorts the picture of both, and it means a Discover fluctuation is investigated as a search problem.

What we do about it

Capabilities

Each of these is work we carry out, not an area we advise on.

Eligibility against the documented requirements

Each published requirement checked against the site's actual output: indexability, the content types Discover surfaces, the image requirements, and the technical conditions that would exclude a page. Establishing this is bounded work with a definite answer.

Image handling and sizing

The image Discover uses, served at the sizes and ratios its requirements specify, present in the delivered HTML, and reachable by the crawler that fetches it. This is where most eligibility failures occur and it is entirely within the site's control.

Delivered markup for feed consumption

What the page provides before JavaScript runs, including the image, the title and the structured data a feed reads. A page whose image arrives through client-side code is a page whose image may not be seen.

Freshness and update signals

Publication and modification times expressed accurately and consistently, since a feed surface depends on knowing when content is new. A modified date that changes on every build is a signal that cannot be used.

Content-type eligibility

Which of the site's content is the kind Discover surfaces, and which is not — so the work is directed at the pages that could be eligible rather than applied across an estate where most of it could not.

Reporting separated from search

Discover measured as its own surface, with its own expectations, so a fluctuation is understood as one rather than investigated as a search regression. This is a reporting change and it prevents a recurring piece of work.

Measurement of a volatile surface

Eligibility, image handling and which content surfaced, rather than traffic totals — because those are the things a site can affect and they are stable enough to be measured. Traffic is reported as what it is: variable by nature.

The boundary, stated

What is within the site's control and what is not, written down. It is the deliverable that stops the engagement becoming an open-ended attempt to influence a system nobody can influence, and it is the most useful thing on the page.

How we work

Engineering methodology

The sequence is deliberate. The order is usually what determines whether the work holds or has to be repeated.

  1. Establish eligibility before attempting anything else

    The documented requirements, checked against the site's output. It has a definite answer, it is bounded, and a page that fails it cannot appear regardless of anything else — so it comes first.

  2. Check the image on the delivered response

    Present in the HTML, at the required size, reachable by a crawler. This is the requirement most often failed and the failure is silent, which is why it is verified rather than assumed from the template.

  3. Direct the work at content that could be eligible

    Discover surfaces particular kinds of content, and applying the requirements across an estate where most pages could never appear is work with no outcome. The content types are established first.

  4. Separate the measurement from search

    Its own surface, its own expectations, its own variability. A publisher who understands that a Discover fluctuation is normal stops investigating it as a regression, which is worth more than most of the technical work.

  5. State the boundary explicitly

    What was made eligible and what cannot be influenced. It is written into the deliverable rather than left implicit, because the alternative is an engagement that continues indefinitely against a system that does not respond to it.

  6. Do not promise inclusion, and do not imply it

    No configuration produces Discover traffic, and a page that reads as though one might is making a claim nobody can support. The work is eligibility and correctness, and the outcome is a page that is not excluded.

Deliverables

What an engagement produces

Documentation is a deliverable, not an afterthought. On most of these engagements a large part of the value is a defect report precise enough for another team to act on.

Eligibility

  • Each documented requirement, checked against the site's actual output
  • The image Discover would use, verified on the delivered response
  • Whether anything excludes the page — robots, rendering, content type
  • Which content types could be eligible, and which could not
  • Publication and modification times, and whether they are usable

Engineering

  • Image handling corrected where it failed a requirement
  • Delivered markup carrying the image, title and structured data
  • Freshness signals accurate and consistent
  • Reporting separated from search

Handover

  • What is eligible now, and what was changed to make it so
  • What remains outside the site's control, stated plainly
  • How to measure this surface without treating its variance as a regression
  • What to check when content changes, so eligibility is not lost
Under the hood

Architecture and technology

What a site controls

  • Indexability, and whether the page can be surfaced at all
  • The image: its presence in the delivered HTML, its size and its ratio
  • Whether the crawler that fetches the image is allowed to
  • The content type, and whether it is one this surface carries
  • Publication and modification times, expressed accurately
  • The structured data a feed reads

What a site does not control

  • Whether any particular page is surfaced to any particular user
  • The volume of traffic, which varies by nature
  • The system that decides what is interesting
  • Its criteria, which are not published
  • When a page stops appearing, or why
Adjacent problems

If this is not quite your problem

These overlap at the edges. Sending you to the right page is more useful than having you work it out.

The site is a publisher

Indexation speed, archive estates and ad-script performance cost.

Publishers

Time-critical content needs its own path

Freshness, eligibility and the news-specific sitemap.

News publisher SEO

The images are the specific problem

Serving the right image at the right size in the delivered markup.

Core Web Vitals optimisation

The content is time-critical and needs a sitemap

A rolling window and per-article freshness.

News sitemap development
Questions

Frequently asked

How do we get more Discover traffic?

No technical change produces Discover traffic, and this is the honest answer rather than a cautious one. Discover surfaces content it judges likely to interest a particular user, from a system whose criteria are not published. What a site can do is meet the documented eligibility requirements and remove anything that would exclude a page — which is bounded work with a definite answer — and then measure honestly. The part that determines whether traffic appears is not the part anyone can configure, and a provider offering Discover traffic is describing something they cannot deliver.

Why did our Discover traffic disappear?

Usually for a reason the site cannot see. The feed changes per user and per session, and traffic moves for reasons unrelated to anything that changed on the site — which is why month-on-month comparison of this surface is close to meaningless and why a disappearance is often not a regression. What is worth checking is whether eligibility was lost: a robots change, a rendering change, an image that stopped being served at the required size. Those are real and checkable. Beyond them, the answer is that the surface varies, and treating its variance as a fault produces work with no outcome.

What image does Discover use?

A large image, and the requirements around its size and ratio are documented. The failure that matters most is not a wrong image but an absent one — an image loaded by client-side code, blocked by a robots rule, or too small to meet the requirement makes a page ineligible, and the failure is not reported anywhere. It presents as a page that simply does not appear. That is why the image is verified on the delivered response rather than assumed from the template.

Do you have Discover-specific work you can point to?

No. Our published engagements are platform and infrastructure engineering, and none of them is Discover-specific, so this page cites nothing and says so. What it describes is the bounded part of the problem: the documented eligibility requirements, image handling, delivered markup and freshness signals — each of which can be checked against your site. The unbounded part is stated as unbounded rather than dressed up as a service.

Should Discover traffic be reported with search?

It arrives in the same reporting and it is a different surface: a feed driven by interest rather than by query, with its own variability. Reporting them together distorts both — a Discover fluctuation looks like a search regression, and it gets investigated as one. Separating them is a small reporting change that prevents a recurring piece of work, and it is usually the most immediately useful outcome of this engagement.

Bring us the problem you have not been able to fix

Describe what is happening rather than what you think the cause is. If we are not the right people for it, we will say so.