Teams/SRE
Site reliability engineers
The engineers who keep your product up. Dedicated SREs working from our Dhaka office, owning your uptime, your monitoring, and your on-call — on one US contract, at flat monthly rates.
prod · reliability
API availability
well within budget
Request latency
comfortable headroom
Background jobs
steady
Alerts
none firing
last incident · timeline
01detect → alert fired on burn rate
02page → on-call acked
03mitigate → rollback complete
04postmortem → drafted, blameless
Downtime is the one bug every customer notices.
Features win customers. Reliability keeps them. A good SRE turns outages from surprises into rare, well-handled events — and turns every incident into a system that fails that way exactly once. That’s the discipline we vet for.
SLOs & error budgets
Reliability targets your team actually agrees on — and a budget that tells you when to ship features and when to stop and fix.
Observability
Metrics, logs, and traces wired in before you need them. When something breaks, you want answers in minutes, not a guessing session.
Incident response
Clear paging, clear roles, calm mitigation. The difference between a bad hour and a bad week is usually process, not luck.
Capacity & scaling
Load tested before launch day, not during it. Autoscaling that's tuned, not just switched on and hoped for.
Blameless postmortems
Every incident ends in a written postmortem that fixes the system, not the person. That's how reliability compounds.
The reliability stack we staff.
Tell us what you run — we match from our pipeline. These are the tools our SREs most commonly own.
| Tooling | What we use it for | Where it shows up |
|---|---|---|
| Prometheus / Grafana | Metrics, dashboards, SLO alerting | The default monitoring stack |
| OpenTelemetry | Traces and instrumentation standards | Distributed systems, microservices |
| PagerDuty-class alerting | On-call schedules and escalation | Any team with a pager |
| Kubernetes | Running and healing production workloads | Most modern infrastructure |
| Load testing — k6, Locust | Finding limits before customers do | Launches, migrations, traffic spikes |
| Chaos / failure testing | Proving systems survive what breaks | Mature reliability programs |
What we look for in an SRE.
Certifications aren't how we vet — our pipeline is. SRE has fewer formal credentials than most disciplines, so we weigh proof of real production ownership. Behind it sits an engineering-school pipeline — BUET, University of Dhaka, NSU and others — and a competitive-programming culture that trains people to stay sharp under pressure.
| Signal | Where it comes from | What it tells us |
|---|---|---|
| CKA — Certified Kubernetes Administrator | CNCF | Can run real workloads on Kubernetes |
| AWS Solutions Architect / SysOps | Amazon | Cloud infrastructure fundamentals |
| Real on-call experience | Previous production roles | Has carried a pager and stayed calm |
| Published postmortems | Past incidents, write-ups | Thinks in systems, writes clearly |
| Monitoring stack ownership | Prior teams | Built observability, didn't just use it |
How we vet a site reliability engineer.
Five stages before you meet anyone — and you still interview last and make the final call. Nobody joins your team without your sign-off.
- 1Incident scenario walkthrough
- 2Systems design under failure
- 3Observability depth interview
- 4Runbook and postmortem review
- 5English & communication screen




Our office in Dhaka — where your team sits. Real photos, real people.
Frequently asked questions
A site reliability engineer keeps production systems available and fast. They set reliability targets (SLOs), build monitoring and observability so problems are caught early, run incident response when something breaks, and write blameless postmortems so the same failure doesn't happen twice. In short, developers ship the product; the SRE makes sure it stays up.
Girmairi's SRE seats are a flat monthly rate: mid-level from $2,800 per month, senior at $4,000, and lead at $4,800, with junior seats available on request. That price includes office, equipment, HR, management, and a replacement guarantee — the full rate card is at girmairi.com/pricing#infrastructure.
Yes. Girmairi's SREs are full-time engineers working from its managed office in Dhaka, Bangladesh — not freelancers — while your contract, NDA, and IP assignment are all with Girmairi LLC, a US company. You get offshore rates with a single US vendor relationship and one monthly USD invoice, billed in advance.
Yes — carrying a pager is part of the job, and real on-call experience is one of the signals SREs are vetted for. On-call coverage and escalation are arranged per engagement, so the schedule fits your timezone, your alerting stack, and how your existing team already handles incidents.
Each SRE passes five stages before you meet anyone: an incident scenario walkthrough, systems design under failure, an observability depth interview, a runbook and postmortem review, and an English and communication screen. You interview last and make the final call — nobody joins your team without your sign-off.
Every seat comes with a replacement guarantee. If an engineer isn't the right fit, you say so and a replacement is sourced from the same vetted pipeline — a bad fit is the vendor's cost, not yours.
Yes. Every client company gets one hour per month with Girmairi's founder — a working CTO — included at no extra cost, useful for reliability strategy, architecture reviews, or pressure-testing an incident process. Committing to 3+ seats or 12 months also earns 5–10% off.
SRE seats are on the standard rate card.
Same flat monthly rate, same includes — office, equipment, HR, management, replacement guarantee.