Search Engineering

Technical SEO and search engineering for high-volume publishers

A newsroom produces content faster than any search engine will fetch it. The engineering question is not how to rank articles — it is how to make sure the ones you published this morning are found while they still have a news life left.

The problem

Publishing volume makes every content-layer assumption fail

Publisher SEO is not the same discipline as ordinary technical SEO at a larger scale. It is a different constraint entirely: content has a short commercially useful life, publication is continuous and cannot be scheduled around technical work, archives are enormous, and the platform is under load precisely when it is producing the most. Advice built for a marketing site — write better titles, improve internal linking — addresses none of it.

  • Articles are indexed days after publication, when the search demand has already passed
  • The news sitemap is not maintained, or not generated from the publishing pipeline
  • Archives are crawled heavily while new articles wait
  • Template-level metadata is wrong across an entire article type at once
  • Crawl rate drops during peak publishing hours, exactly when it matters
  • Recurring server errors interrupt crawls at predictable times
  • Category and tag pages compete with the articles they list
  • Editorial and engineering teams are not working from the same set of facts
Who this is for

The people who usually bring us this problem

Head of SEO at a publisher

You know discovery is slower than it should be, you can see it in the crawl and index data, and the fix is in the publishing pipeline rather than in editorial.

Head of Audience / Editorial Operations

Your journalists are producing and the audience is not arriving in the first hours. You need to know whether that is a content problem or a delivery one.

CTO / Platform Lead at a media business

Crawl behaviour, archive scaling and publishing peaks are infrastructure concerns presenting as search concerns, and you need them described in engineering terms.

What it costs

What this costs while it goes unfixed

Engineering faults are rarely confined to the engineering layer. These are the commercial consequences we see most often.

News content loses most of its value if discovery is slow

An article found four days after publication has already missed the demand it was commissioned to meet. This is not a ranking penalty; it is the news cycle moving on, and nothing recovers it afterwards.

Crawl budget goes to the archive instead of the newsroom

A large archive is an easy crawl target and a new article is not. Without deliberate prioritisation, a search engine spends its visit on pages whose commercial life is over while today's output waits.

A template defect affects an entire article type at once

Publishers run a small number of templates across a very large number of pages. A metadata or canonical fault in one template suppresses every article using it, and the symptom looks like declining content quality rather than a technical fault.

Server reliability and discovery are the same problem

Crawlers that meet errors during a visit reduce their frequency. Recurring 5xx at predictable times is one of the most direct ways to lose discovery speed, and it presents as an SEO problem rather than an infrastructure one.

What we do about it

Capabilities

Each of these is work we carry out, not an area we advise on.

Discovery speed engineering

The whole path from publish to discoverable: the event that fires on publication, the feeds that announce it, the freshness signals in the markup, and whether the search engine's next visit finds the article ready. This is the core of the engagement and most of its value.

News sitemap architecture

A sitemap generated from the publishing pipeline rather than exported on a schedule, correct publication timestamps, eligibility filtering, and freshness monitoring so a stalled generator is noticed.

News sitemap development in depth

Article and archive templating

Making the template layer produce correct metadata, canonicals, structured data and pagination as a property of the system, so no article depends on an editor remembering to set something.

Archive and pagination strategy

Deciding deliberately what the archive is for: what is indexable, how paginated sequences signal themselves, how dated and expired content is treated, and how much crawl attention the archive is allowed to claim.

Crawl efficiency

Reducing wasted crawl on category, tag, author and faceted URLs so the available crawl budget reaches articles. On a large publisher this is frequently a bigger lever than anything applied to the articles themselves.

Server reliability as a discovery factor

Diagnosing the error and latency patterns that reduce crawl frequency, and coordinating their resolution with the infrastructure work — because on a publishing platform, availability during production hours is a search consideration.

Structured data at article scale

Article-specific markup emitted from the publishing system consistently across every article type, rather than configured per page or applied to the most recent content only.

Measurement that editorial and engineering share

Instrumentation that shows discovery and publication together, so the two teams can work from one set of facts rather than from two conflicting interpretations of the same decline.

How we work

Engineering methodology

The sequence is deliberate. The order is usually what determines whether the work holds or has to be repeated.

  1. Start with the publish event, end with the crawl

    The full path from a story going live to a search engine fetching it is traced end to end. Discovery failures are usually a break at one identifiable point along it, and the fastest way to find that point is to walk the path rather than to begin with a list of recommendations.

  2. Separate the four constraints, because they have four fixes

    Discovery speed, crawl allocation, template correctness and server reliability produce similar-looking symptoms and are resolved by different work. Diagnosing which is binding is the first deliverable, and it is frequently not the one the client expected.

  3. Fix at the template layer, not the page layer

    On a publisher, a correction applied to one article is a correction applied to nothing. Every fix is expressed as a change to the template or the pipeline that produces the pages, so it applies to content that does not exist yet.

  4. Protect the news window before optimising anything historical

    Improving discovery of today's output has a far higher return than improving the performance of a two-year-old archive. Work is prioritised by remaining commercial life, which is a different ordering from the one conventional SEO recommendations arrive in.

  5. Coordinate with the teams that own the platform

    Publisher SEO work reaches into the CMS, the publishing pipeline and the infrastructure. It is sequenced with the people who own those, and the coordination is part of the engagement rather than a hand-off at the end.

  6. Measure discovery against publication, over time

    Time from publish to first crawl and to first impression, tracked continuously rather than assessed once. Publisher SEO is a property of an operating system, not a project with a completion date.

Deliverables

What an engagement produces

Documentation is a deliverable, not an afterthought. On most of these engagements a large part of the value is a defect report precise enough for another team to act on.

Diagnosis

  • The publish-to-crawl path traced end to end, with the break identified
  • Which of discovery speed, crawl allocation, template correctness or reliability is binding
  • Crawl distribution: archive, category and article share of crawl attention
  • Template-level defect review across every article type
  • Measurement framework agreed with both editorial and engineering

Engineering

  • Discovery path fixes: publish event, feed generation and freshness signals
  • News sitemap generated from the publishing pipeline
  • Article template correctness at the system level
  • Archive and pagination strategy implemented
  • Crawl waste reduced on category, tag, author and faceted URLs
  • Structured data emitted consistently across article types
  • Server error coordination where reliability is suppressing crawl

Operating

  • Discovery measured against publication on an ongoing basis
  • Freshness monitoring on the feeds
  • Template regression checking, so a redesign cannot silently suppress a page type
  • A shared dashboard that editorial and engineering both use
Under the hood

Architecture and technology

What a publisher's discovery path consists of

  • The publish event, from the CMS or publishing service that owns the article lifecycle
  • Feed generation, including the news sitemap and any other announcement channel
  • The article template, emitting metadata, canonicals and structured data
  • Freshness signals that agree with the actual publication time
  • Crawl allocation between articles, archives and taxonomy pages
  • Server availability during the hours when publishing and crawling overlap
  • Measurement connecting publication to discovery

Publisher-specific constraints

  • Publication is continuous and cannot pause for technical work
  • Content has a short commercially useful life, which reorders every priority
  • A small number of templates cover a very large number of pages
  • Archives are large, crawl-attractive and of declining value
  • Load peaks coincide with production peaks
  • Editorial and engineering ownership are usually separate
Related work

Where we have done this

Engagements where this capability was the substance of the work rather than a line item.

Fintech & capital markets

Search engineering at stockbroking scale

A stockbroking platform publishing at news velocity was losing search visibility to problems that had nothing to do with content quality. Two workstreams ran in parallel: sustaining a high-volume editorial output across business and market categories, and diagnosing the technical faults — a domain safety flag, recurring server errors, and a metadata defect on a templated page type — that were suppressing how much of that output search engines could actually reach.

225Msitewide impressions
Adjacent problems

If this is not quite your problem

These overlap at the edges. Sending you to the right page is more useful than having you work it out.

The specific problem is the news feed

Feed architecture, validation and freshness monitoring in depth.

News sitemap development

Articles are crawled but not indexed

Discovery has succeeded and something downstream has not.

Crawl and indexation diagnosis

Server errors are interrupting crawls

Availability is an engineering problem before it is a search one.

Platform reliability

The context is publishing as a sector

Publisher constraints, business models and platform patterns.

Publishers and news
Questions

Frequently asked

Is publisher SEO different from technical SEO?

Yes, and the difference is not scale. Technical SEO asks how a site can be crawled and indexed; publisher SEO asks how output can be discovered while it still has commercial life, which is a question about latency rather than about reachability. It reorders the priorities entirely — fixing discovery for today's articles matters far more than improving the performance of an archive that is already past its value.

Will a news sitemap fix our indexing?

It addresses discovery, which is one of four constraints and often the most valuable one — but it does not force indexing, and a sitemap cannot compensate for a template defect or a server that fails during crawls. If your articles are being crawled promptly and not indexed, the sitemap is not the problem and improving it will not help. Establishing which constraint is binding comes before recommending a fix.

How quickly will we see results?

Discovery improvements can show up within weeks, because they change how quickly a search engine learns about new articles — which is a shorter feedback loop than ranking changes. Anything involving crawl allocation takes longer, because it requires the search engine to change its behaviour toward your site over many visits. We track time from publication to first crawl and to first impression, so the effect is visible as a trend rather than inferred.

Can you guarantee our traffic will increase?

No, and a publisher is the worst possible context in which to promise it. Discovery is a technical property we can improve and measure; whether an article earns demand depends on the story, the market and the news cycle, none of which is ours. What we commit to is a faster and more reliable path from publication to being findable, verified against measurement.

Do you work with our editorial team or only engineering?

Both, because on a publisher the two own different halves of the same outcome. Engineering owns the pipeline, the templates and the infrastructure; editorial owns the publication cadence and the content types. Most publisher discovery problems we are asked about turn out to sit at the boundary between them, which is why the measurement framework is designed for both teams to read rather than for one to report to the other.

Our traffic dropped after a redesign. Where do we start?

With the template layer, because a redesign changes the templates that generate every article and a single defect there affects the whole page type at once. The first thing to establish is whether the drop is uniform across article types or concentrated in one — that distinction usually identifies whether the cause is a template, a crawl change or something outside the site entirely.

Bring us the problem you have not been able to fix

Describe what is happening rather than what you think the cause is. If we are not the right people for it, we will say so.