Engineering

Infrastructure engineering and managed operations

The layer that decides whether everything above it works. We design, provision, tune and then operate the infrastructure a platform runs on — including the ongoing management that keeps a tuned system tuned.

The problem

Infrastructure is where reliability is decided

Application performance, availability, security posture and cost all resolve to infrastructure decisions that are usually made once, early, and then inherited. A platform running on infrastructure designed for a much smaller version of itself will exhibit symptoms that look like application problems and are not.

  • Resource utilisation that looks fine while users experience slowness
  • Shared hosting or a general-purpose environment under production traffic
  • No clear answer to which resource will saturate first
  • No visibility into server behaviour until a user reports a problem
  • Infrastructure cost growing without a corresponding change in capability
  • Configuration that has never been reviewed since it was first set up
Who this is for

The people who usually bring us this problem

CTO / VP Engineering

You need the infrastructure layer owned, tuned and monitored rather than treated as an inherited constant.

Founder

You are running a platform whose reliability matters commercially, without an infrastructure engineer to own it.

Head of Product

Reliability problems are surfacing as user complaints, and you need the underlying cause identified.

What it costs

What this costs while it goes unfixed

Engineering faults are rarely confined to the engineering layer. These are the commercial consequences we see most often.

Undersized infrastructure constrains everything above it

Application optimisation cannot overcome a platform that has run out of the resource it needs. Symptoms present as application faults and lead teams to optimise the wrong tier.

Unmanaged infrastructure drifts

A server that was correctly tuned two years ago is not correctly tuned now: traffic, software versions and data volumes have all changed. Without review, configuration becomes progressively less appropriate.

Simplified hosting is cheap until it is not

Shared and general-purpose environments are economical at small scale and become a ceiling at larger scale, frequently at the least convenient moment.

What we do about it

Capabilities

Each of these is work we carry out, not an area we advise on.

Infrastructure assessment

Establishing what the platform actually needs — by resource, by concurrency and by growth trajectory — rather than what it currently has.

Server provisioning and configuration

Dedicated and VPS environments sized and configured for the observed load profile, including the concurrency settings that default configurations rarely get right.

CDN and edge architecture

Designing what is served from where, and with what cache policy. Getting this right removes a large share of load from the origin entirely.

Load balancing and high availability

Distributing demand and removing single points of failure, so that capacity and availability are not the same resource.

Caching architecture

Layered caching where the correctness of the invalidation policy matters as much as the hit rate.

Monitoring and observability

Instrumentation that surfaces developing problems before they become incidents, with alerting calibrated to be actionable rather than ignored.

Backup, recovery and business continuity

Backups that have been tested by restoring from them, and a recovery procedure that has been rehearsed rather than assumed.

Managed operations

Ongoing management of the infrastructure estate: patching, capacity review, incident response and continuous tuning as the platform changes.

How we work

Engineering methodology

The sequence is deliberate. The order is usually what determines whether the work holds or has to be repeated.

  1. Measure before sizing

    Capacity decisions are made from observed load profiles, including peaks, not from server specifications or rules of thumb.

  2. Remove load before adding capacity

    Caching and edge delivery are almost always cheaper and more durable than provisioning to handle work that did not need to reach the origin.

  3. Design for the peak, operate for the average

    Sizing must accommodate the worst case; day-to-day operation is tuned for the common case. Conflating the two produces either fragile systems or wasteful ones.

  4. Make it observable before trusting it

    Infrastructure without instrumentation is infrastructure whose behaviour is unknown. Monitoring is delivered as part of the build, not as an addition.

  5. Then keep managing it

    A tuned system does not stay tuned. Continuous management is what converts an infrastructure project into sustained reliability.

Deliverables

What an engagement produces

Documentation is a deliverable, not an afterthought. On most of these engagements a large part of the value is a defect report precise enough for another team to act on.

Assessment

  • Current infrastructure and load analysis
  • Resource-constraint identification
  • Capacity and growth modelling
  • Cost analysis against capability

Build

  • Provisioning and configuration
  • CDN and edge architecture
  • Load balancing and redundancy
  • Caching layers with correct invalidation

Operate

  • Monitoring and alerting
  • Tested backup and recovery
  • Patching and version management
  • Capacity review and continuous tuning
  • Incident response
Under the hood

Architecture and technology

Layers

  • Edge and CDN
  • Load balancing and routing
  • Application servers
  • Cache layers
  • Database and storage
  • Backup and recovery

Observability

  • Resource utilisation and saturation
  • Response-time percentiles rather than averages
  • Error rate and distribution
  • Capacity headroom and trend
Related work

Where we have done this

Engagements where this capability was the substance of the work rather than a line item.

Online communities

Spain's largest forum

Since 2019, Soludome has managed infrastructure and platform operations for one of Spain's largest online communities. The work covers server and performance optimisation, CDN architecture, load balancing and site security under sustained traffic.

Since 2019continuous engagement
Fintech & capital markets

Search engineering at stockbroking scale

A stockbroking platform publishing at news velocity was losing search visibility to problems that had nothing to do with content quality. Two workstreams ran in parallel: sustaining a high-volume editorial output across business and market categories, and diagnosing the technical faults — a domain safety flag, recurring server errors, and a metadata defect on a templated page type — that were suppressing how much of that output search engines could actually reach.

225Msitewide impressions
Marketplaces

Platform recovery

A large marketplace platform in the United States was experiencing critical downtime from compounding traffic and application faults. Soludome analysed requests through the web application firewall, identified and blocked the DDoS patterns, and repaired a third-party payment flow that allowed links to be generated repeatedly without authentication. Server load fell and platform performance improved after the work.

Around 2020initial recovery; occasional support since
Online communities & specialist publishing

GuzziTech & RideMalibu

Since 2021, Soludome has managed infrastructure, site performance and technical SEO across GuzziTech and RideMalibu as a connected web estate. The engagement includes dedicated infrastructure and continuing operational ownership.

Since 2021managed infrastructure, performance and SEO
Adjacent problems

If this is not quite your problem

These overlap at the edges. Sending you to the right page is more useful than having you work it out.

The infrastructure is fine but pages are slow

The constraint is in the page rather than beneath it. See Website Performance Engineering.

The platform fails specifically under load

That is a concurrency and architecture problem. See High-Traffic Website Engineering.

Questions

Frequently asked

Do you provide hosting?

We provision and manage infrastructure, sometimes on the client's own accounts and sometimes on ours, depending on what suits. Where we manage it, the client retains ownership and access — infrastructure a client cannot inspect is a dependency rather than a service.

We are on shared hosting. Is that a problem?

It depends entirely on your load profile, and it is a question of when rather than whether. Shared environments are economical and perfectly adequate at small scale. The problem is that the ceiling arrives with growth, and the migration is far less disruptive if planned than if forced by an outage.

What does ongoing management actually cover?

Monitoring and alerting, patching and version management, capacity review against growth, incident response, and continuous tuning as the platform changes. The point is that infrastructure is not a state you reach but a thing you maintain — most of the value in this practice is in the management rather than the initial build.

How do you handle backups?

By testing restoration rather than the backup. An untested backup is a belief, not a control. We deliver backups that have been restored from, with a recovery procedure that has been rehearsed and a documented recovery time.

Bring us the problem you have not been able to fix

Describe what is happening rather than what you think the cause is. If we are not the right people for it, we will say so.