Platform Engineering

Managed server services for reliable websites and applications

Somebody has to own the server. If it is nobody in particular, it is owned by whoever notices a problem last. We take responsibility for agreed operations — patching, backups, monitoring and escalation — with the boundary of that responsibility written down before the first change.

The problem

The server is everyone's problem and nobody's job

A production environment usually drifts into being unmanaged rather than being left unmanaged deliberately. It was set up competently by someone who has since moved on, it works, and it is nobody's assigned responsibility until it fails. The failures that follow are rarely dramatic: an untested backup, a patch that was never applied, a certificate that expired on a Sunday, a disk that filled quietly for six weeks.

  • Nobody can say confidently who is responsible for server updates
  • Backups run, but nobody has restored one to prove they work
  • Everything runs on one machine with no separation between application and database
  • Patching happens reactively, after something has been exploited
  • Capacity problems are discovered when the site goes down rather than in advance
  • Alerts arrive somewhere nobody is looking, or nowhere at all
  • The person who set the server up has left and there is no documentation
  • A deployment procedure exists in someone's head rather than in writing
Who this is for

The people who usually bring us this problem

A business with a site but no infrastructure owner

Your website matters to revenue and its server has no named owner. You want that responsibility held formally rather than by whoever happens to notice.

CTO / VP Engineering with a small team

Your engineers should be building the product, not chasing certificate renewals and disk usage. You need operations handled with clear escalation and a written boundary.

Head of Digital / Operations

You have inherited a production environment with no documentation and you need a competent inventory before you can even decide what to change.

What it costs

What this costs while it goes unfixed

Engineering faults are rarely confined to the engineering layer. These are the commercial consequences we see most often.

Backups nobody has restored are a hypothesis, not a backup

A backup routine that has never been tested is a routine, not a recovery capability. The failure is discovered at the worst possible moment, and it is completely avoidable.

Deferred patching compounds into a forced migration

Each version skipped makes the next one harder. Two years of deferral turns a routine update into a project with a deadline set by someone else, usually because support has ended.

Reactive operations cost more than the retainer they replace

An outage is paid for in lost revenue, emergency work at unsocial hours and, frequently, an improvised fix that becomes the next problem. The recurring cost of avoiding it is almost always lower.

Undocumented environments cannot be handed over

When knowledge lives in one person's head, their departure is an incident. Documentation is not administration overhead; it is what makes the environment survivable.

What we do about it

Capabilities

Each of these is work we carry out, not an area we advise on.

Onboarding inventory

Before anything is changed, a written record of what exists: services, versions, scheduled jobs, integrations, DNS, certificates, backup destinations and who holds which access. Most of this does not currently exist for the environments we are asked to take on, and it is the first deliverable rather than an afterthought.

Access model

Establishing who can reach the environment, through what, and with what privilege — replacing shared credentials and undocumented keys with named access that can be revoked when someone leaves.

Patch and version management

A defined schedule rather than a policy: what is updated, how it is tested, when it is applied, and what the rollback is if it goes wrong.

Backup verification

Confirming that backups exist, that they complete, and — the part that is usually missing — that a restore actually works. Tested restore is the deliverable; a backup log is not evidence.

Monitoring and alerting

Availability, resource use, certificate expiry, disk growth and service health, wired to somewhere a human actually reads. Alerting nobody receives is indistinguishable from no alerting.

Capacity and performance

Web server, database and application-level tuning reviewed against observed load, with headroom estimated in advance rather than discovered during an incident.

Deployment and change procedure

A documented, repeatable way to ship a change — including how to reverse it — so releases stop being unrepeatable events whose success depends on the person doing them.

Escalation path

A named route for a problem that exceeds the agreed scope, rather than an informal expectation that someone will notice a message.

Written reporting

A recurring report of what was done, what was found and what needs attention, so the state of the environment is visible to the people funding it rather than inferred from the absence of complaints.

How we work

Engineering methodology

The sequence is deliberate. The order is usually what determines whether the work holds or has to be repeated.

  1. Inventory before intervention

    We do not change a production environment we have not first described. The inventory is written down and reviewed with you, because half of the value of taking on an undocumented server is producing the document nobody had.

  2. Define the boundary in writing first

    A shared-responsibility model is agreed before work starts: what we operate, what remains yours, what the hosting provider owns, and where each ends. Most managed-server disputes are disagreements about this table that were never made explicit.

  3. Verify the recovery path before optimising anything

    A tested restore is established early, because until it exists every other change carries avoidable risk. This is usually the least glamorous part of onboarding and the most valuable.

  4. Make monitoring observable, then act on it

    Alerting is set up so that it reaches a human, and then the alerts are actually reviewed. Monitoring that produces noise nobody reads is worse than no monitoring, because it creates the impression of coverage.

  5. Change on a schedule, not under pressure

    Patching, updates and version upgrades run on a defined cadence with a tested rollback. Emergency changes are occasionally necessary and are always a signal that something upstream was deferred.

  6. Report so the environment is legible

    A recurring written report means the state of the platform is a fact the business can see, rather than a subject that only comes up during an outage.

Deliverables

What an engagement produces

Documentation is a deliverable, not an afterthought. On most of these engagements a large part of the value is a defect report precise enough for another team to act on.

Onboarding

  • Environment inventory: services, versions, jobs and dependencies
  • Access review and consolidation
  • Backup verification including a tested restore
  • Monitoring baseline and alert routing
  • Agreed shared-responsibility matrix
  • Documented escalation path and scope exclusions

Operating

  • Scheduled patching and version updates with rollback
  • Availability and resource monitoring
  • Certificate and DNS renewal tracked in advance
  • Backup monitoring and periodic restore testing
  • Performance and capacity review against observed load
  • Deployment procedure maintained and executed

Reporting and handover

  • Recurring written report of work, findings and risks
  • Change log attributed and dated
  • Documentation kept current as the environment changes
  • An explicit handover pack if the engagement ends
Under the hood

Architecture and technology

Shared responsibility, stated plainly

  • Ours: the server, its configuration, its patching, its backups and its availability
  • Yours: the application, its code, its content and its release decisions
  • The hosting provider's: the physical or virtual layer, the network and the datacentre
  • Joint, by agreement: deployment procedure, capacity decisions and incident response
  • Documented once, reviewed when anything material changes

What we do not claim

  • We do not operate a datacentre and we do not resell hosting
  • We do not advertise a control panel or operating-system range we have not verified we can support
  • We do not provide round-the-clock cover unless it is separately contracted
  • We do not take on an environment we cannot inventory
Related work

Where we have done this

Engagements where this capability was the substance of the work rather than a line item.

Online communities

Spain's largest forum

Since 2019, Soludome has managed infrastructure and platform operations for one of Spain's largest online communities. The work covers server and performance optimisation, CDN architecture, load balancing and site security under sustained traffic.

Since 2019continuous engagement
Online communities & specialist publishing

GuzziTech & RideMalibu

Since 2021, Soludome has managed infrastructure, site performance and technical SEO across GuzziTech and RideMalibu as a connected web estate. The engagement includes dedicated infrastructure and continuing operational ownership.

Since 2021managed infrastructure, performance and SEO
Adjacent problems

If this is not quite your problem

These overlap at the edges. Sending you to the right page is more useful than having you work it out.

You are moving to a new server rather than running one

The cutover, DNS and validation work is a distinct engagement.

Server migration with a tested cutover

The server is failing intermittently rather than being unmanaged

Recurring outages and 5xx errors are a diagnosis problem before they are an operations problem.

Platform reliability

You need the architecture reviewed rather than operated

Capacity design, topology and scaling decisions are a different brief from ongoing operations.

Infrastructure engineering

You want a platform managed rather than a server

XenForo communities under managed infrastructure are covered separately.

Managed XenForo hosting
Questions

Frequently asked

Can you manage an unmanaged VPS?

Yes, and that is most of what we take on. Onboarding starts from whatever state the environment is in — including no documentation at all — and the first deliverable is an inventory of what actually exists. We would rather describe the environment accurately than assume it matches how it was originally specified.

Which operating systems and control panels do you support?

We are deliberately not listing them here. A support matrix that has not been verified against what we can actually operate is the easiest kind of untrue claim to publish, and the cost of getting it wrong lands on you at the point you need help. We will tell you plainly, for your specific environment, what we can take responsibility for and what we cannot.

Do you take over an existing server from another provider?

Yes. We start with an inventory and an access review, then establish a tested restore path before changing anything. The most common finding is that the previous setup is broadly sound but undocumented and untested — which is a different problem from a bad configuration, and cheaper to fix.

What are your response and coverage hours?

They are set per engagement rather than published as a standard, because a practice of this size cannot honestly staff a round-the-clock guarantee. What is agreed is written into the shared-responsibility matrix alongside the escalation path, so there is no ambiguity at the moment you need it — which is the only moment it matters.

Do you resell hosting or operate your own datacentre?

No to both. We operate environments on infrastructure you or your provider own, which means you keep the commercial relationship and the ability to move. We do not mark up hosting, and we are not a reseller of anybody's platform.

What happens if we stop working together?

You receive a handover pack: the current inventory, the change log, the runbooks, and the credentials returned to your control. The environment should be in a state another competent operator can pick up, because everything we did was documented as we went rather than reconstructed at the end.

Bring us the problem you have not been able to fix

Describe what is happening rather than what you think the cause is. If we are not the right people for it, we will say so.