Releases feel risky
Deploys are manual, rollback is written down nowhere, and it all rests on one person. The question is not what is wrong, it is where to start.
Fractional Platform / SRE Principal
I work out where infrastructure and backend create risk for the product, and help you decide what to fix first. I do the work myself, and it stays with your team afterwards.
For product companies with in-house engineering that already have the platform work, but not yet the senior engineer to own it.
When to bring me in
Deploys are manual, rollback is written down nowhere, and it all rests on one person. The question is not what is wrong, it is where to start.
Backups run and alerts fire. Whether recovery actually works tends to get answered at the worst possible moment.
Infrastructure is getting complicated faster than the team is growing. Worth reviewing boundaries, dependencies and how it is operated before a migration or the next stage.
Services
Choose one question and an agreed sample of systems. Scope, exclusions, deliverables, dates and price are agreed before work begins.
Where do infrastructure and the delivery process create risk for the product?
First engagementReport, executive summary and a team review
Which failure scenarios affect users, and what should you address first?
First engagementReliability priorities with owners and dependencies
Implement one selected improvement after agreeing the priorities.
Next stepA fixed price for the agreed deliverable
Assessments use read-only access. Implementation and active tests are agreed separately. If the need turns out to be recurring, we agree an ongoing arrangement with defined monthly capacity. I do not offer 24/7 cover under any of them.
Experience
Over 13 years in development and operations: Go and Python backend, platforms, reliability, infrastructure in the cloud and on bare metal. I handle the conversation, the work and the handover myself — no subcontractors, no juniors on your project.
Kubernetes, IaC and GitOps, CI/CD, PostgreSQL, queues, observability and recovery — selected to fit the problem and the team's operating capacity.
Device traffic in Go, rentals and payments in Python, an event queue between them. Plus a gateway stand-in, so connection faults get caught off site.
How it worksReliability built as rungs: a second copy of the database, connection pooling and logs arrive as separate layers, switched on when the system grows into them.
How it worksThe whole estate described in one variable, with the Ansible host list rendered from it. Copying addresses by hand is gone.
How it worksThese are my own products and internal automation. Client work under NDA is not shown here.
How engagements work
Use 30 minutes to clarify the problem, urgency, owner and constraints. Agree one next step.
Define the sample, exclusions, deliverables, price and acceptance criteria. Estimate effort separately from calendar dates.
Once start conditions are met, confirm roles, access and baseline evidence. Keep unknowns explicit.
Collect evidence and discuss risks and decisions. Receive a weekly status update; additional work is agreed separately.
Review the report and roadmap, transfer the materials, record acceptance and revoke access. Implementation is a separate engagement.
One main project at a time. Price and dates depend on the agreed scope and team availability. Contract and payment arrangements are confirmed for each engagement.
Contact
Tell me what the product does, what is slowing the team down and why this matters now. Please leave out secrets, credentials and user data.
The first conversation is 30 minutes to clarify the task. Technical assessment and architecture recommendations belong to the agreed paid engagement.
You leave with one concrete next step: a scoped assessment, an implementation sprint, or advice to come back to this later. The third one comes up more often than people admit.