Fractional Platform / SRE Principal

Production reliability and a clear plan for change

I work out where infrastructure and backend create risk for the product, and help you decide what to fix first. I do the work myself, and it stays with your team afterwards.

For product companies with in-house engineering that already have the platform work, but not yet the senior engineer to own it.

Diagram: client reaches the API, the primary database fails, and the path leads to a standby database
Review dependencies, failure scenarios and recovery paths.

When to bring me in

Releases, recovery and growth need attention

Releases feel risky

Deploys are manual, rollback is written down nowhere, and it all rests on one person. The question is not what is wrong, it is where to start.

Incidents keep recurring

Backups run and alerts fire. Whether recovery actually works tends to get answered at the worst possible moment.

The team is growing

Infrastructure is getting complicated faster than the team is growing. Worth reviewing boundaries, dependencies and how it is operated before a migration or the next stage.

Services

Start with a bounded paid assessment

Choose one question and an agreed sample of systems. Scope, exclusions, deliverables, dates and price are agreed before work begins.

Platform model under a magnifying glass

Infrastructure / Platform Health Check

Where do infrastructure and the delivery process create risk for the product?

  • Current architecture and baseline
  • Selected CI/CD, IaC, Kubernetes and observability checks
  • Findings with evidence and limitations
  • Priorities and a 30/60/90-day roadmap

First engagementReport, executive summary and a team review

A bypass bridge connects modules around a failed component

Reliability Assessment

Which failure scenarios affect users, and what should you address first?

  • One critical user journey
  • Incident history, alerts and ownership
  • SLIs/SLOs and recovery evidence
  • Risks, unknowns and a measurement plan

First engagementReliability priorities with owners and dependencies

A new module closes the gap in a platform; the old one has been replaced

Implementation Sprint

Implement one selected improvement after agreeing the priorities.

  • A separate scope and acceptance criteria
  • For example: rollback, a recovery runbook or a release path
  • Verification in an agreed environment
  • Documentation and knowledge transfer

Next stepA fixed price for the agreed deliverable

Assessments use read-only access. Implementation and active tests are agreed separately. If the need turns out to be recurring, we agree an ongoing arrangement with defined monthly capacity. I do not offer 24/7 cover under any of them.

Experience

Aleksandr Sapon — the engineer delivering your project

Over 13 years in development and operations: Go and Python backend, platforms, reliability, infrastructure in the cloud and on bare metal. I handle the conversation, the work and the handover myself — no subcontractors, no juniors on your project.

Kubernetes, IaC and GitOps, CI/CD, PostgreSQL, queues, observability and recovery — selected to fit the problem and the team's operating capacity.

Own product · 2026

Tool rental through smart lockers

Device traffic in Go, rentals and payments in Python, an event queue between them. Plus a gateway stand-in, so connection faults get caught off site.

How it works
Own product · 2026

White-label rental platform

Reliability built as rungs: a second copy of the database, connection pooling and logs arrive as separate layers, switched on when the system grows into them.

How it works
Internal automation · 2020

Environments for development teams

The whole estate described in one variable, with the Ansible host list rendered from it. Copying addresses by hand is gone.

How it works

These are my own products and internal automation. Client work under NDA is not shown here.

How engagements work

From a question to a report and the next decision

01

Conversation

Use 30 minutes to clarify the problem, urgency, owner and constraints. Agree one next step.

02

Proposal and SOW

Define the sample, exclusions, deliverables, price and acceptance criteria. Estimate effort separately from calendar dates.

03

Kickoff and baseline

Once start conditions are met, confirm roles, access and baseline evidence. Keep unknowns explicit.

04

Review and report

Collect evidence and discuss risks and decisions. Receive a weekly status update; additional work is agreed separately.

05

Handover

Review the report and roadmap, transfer the materials, record acceptance and revoke access. Implementation is a separate engagement.

One main project at a time. Price and dates depend on the agreed scope and team availability. Contract and payment arrangements are confirmed for each engagement.

Contact

Start with your situation

Tell me what the product does, what is slowing the team down and why this matters now. Please leave out secrets, credentials and user data.

The first conversation is 30 minutes to clarify the task. Technical assessment and architecture recommendations belong to the agreed paid engagement.

You leave with one concrete next step: a scoped assessment, an implementation sprint, or advice to come back to this later. The third one comes up more often than people admit.