Girmairi for Saudi Arabia:Read the Vision 2030 brief ↗
TeamsSolutionsResourcesCompanyBlogRate cardFAQ
Get in touch

Teams/SRE

Site reliability engineers

The engineers who keep your product up. Dedicated SREs working from our Dhaka office, owning your uptime, your monitoring, and your on-call — on one US contract, at flat monthly rates.

SLOsObservabilityIncident responseKubernetesPrometheusOn-call

prod · reliability

API availability

well within budget

Request latency

comfortable headroom

Background jobs

steady

Alerts

none firing

last incident · timeline

01detect → alert fired on burn rate

02page → on-call acked

03mitigate → rollback complete

04postmortem → drafted, blameless

Why SRE

Downtime is the one bug every customer notices.

Features win customers. Reliability keeps them. A good SRE turns outages from surprises into rare, well-handled events — and turns every incident into a system that fails that way exactly once. That’s the discipline we vet for.

01

SLOs & error budgets

Reliability targets your team actually agrees on — and a budget that tells you when to ship features and when to stop and fix.

02

Observability

Metrics, logs, and traces wired in before you need them. When something breaks, you want answers in minutes, not a guessing session.

03

Incident response

Clear paging, clear roles, calm mitigation. The difference between a bad hour and a bad week is usually process, not luck.

04

Capacity & scaling

Load tested before launch day, not during it. Autoscaling that's tuned, not just switched on and hoped for.

05

Blameless postmortems

Every incident ends in a written postmortem that fixes the system, not the person. That's how reliability compounds.

Tooling

The reliability stack we staff.

Tell us what you run — we match from our pipeline. These are the tools our SREs most commonly own.

ToolingWhat we use it forWhere it shows up
Prometheus / GrafanaMetrics, dashboards, SLO alertingThe default monitoring stack
OpenTelemetryTraces and instrumentation standardsDistributed systems, microservices
PagerDuty-class alertingOn-call schedules and escalationAny team with a pager
KubernetesRunning and healing production workloadsMost modern infrastructure
Load testing — k6, LocustFinding limits before customers doLaunches, migrations, traffic spikes
Chaos / failure testingProving systems survive what breaksMature reliability programs
Credentials & signals

What we look for in an SRE.

Certifications aren't how we vet — our pipeline is. SRE has fewer formal credentials than most disciplines, so we weigh proof of real production ownership. Behind it sits an engineering-school pipeline — BUET, University of Dhaka, NSU and others — and a competitive-programming culture that trains people to stay sharp under pressure.

SignalWhere it comes fromWhat it tells us
CKA — Certified Kubernetes AdministratorCNCFCan run real workloads on Kubernetes
AWS Solutions Architect / SysOpsAmazonCloud infrastructure fundamentals
Real on-call experiencePrevious production rolesHas carried a pager and stayed calm
Published postmortemsPast incidents, write-upsThinks in systems, writes clearly
Monitoring stack ownershipPrior teamsBuilt observability, didn't just use it
Vetting

How we vet a site reliability engineer.

Five stages before you meet anyone — and you still interview last and make the final call. Nobody joins your team without your sign-off.

  1. 1Incident scenario walkthrough
  2. 2Systems design under failure
  3. 3Observability depth interview
  4. 4Runbook and postmortem review
  5. 5English & communication screen
Engineers pairing at a desk in Girmairi's Dhaka office
Meeting corner in Girmairi's Dhaka office
The engineering floor at Girmairi's Dhaka office
Training hall in Girmairi's Dhaka office

Our office in Dhaka — where your team sits. Real photos, real people.

Frequently asked questions

A site reliability engineer keeps production systems available and fast. They set reliability targets (SLOs), build monitoring and observability so problems are caught early, run incident response when something breaks, and write blameless postmortems so the same failure doesn't happen twice. In short, developers ship the product; the SRE makes sure it stays up.

Girmairi's SRE seats are a flat monthly rate: mid-level from $2,800 per month, senior at $4,000, and lead at $4,800, with junior seats available on request. That price includes office, equipment, HR, management, and a replacement guarantee — the full rate card is at girmairi.com/pricing#infrastructure.

Yes. Girmairi's SREs are full-time engineers working from its managed office in Dhaka, Bangladesh — not freelancers — while your contract, NDA, and IP assignment are all with Girmairi LLC, a US company. You get offshore rates with a single US vendor relationship and one monthly USD invoice, billed in advance.

Yes — carrying a pager is part of the job, and real on-call experience is one of the signals SREs are vetted for. On-call coverage and escalation are arranged per engagement, so the schedule fits your timezone, your alerting stack, and how your existing team already handles incidents.

Each SRE passes five stages before you meet anyone: an incident scenario walkthrough, systems design under failure, an observability depth interview, a runbook and postmortem review, and an English and communication screen. You interview last and make the final call — nobody joins your team without your sign-off.

Every seat comes with a replacement guarantee. If an engineer isn't the right fit, you say so and a replacement is sourced from the same vetted pipeline — a bad fit is the vendor's cost, not yours.

Yes. Every client company gets one hour per month with Girmairi's founder — a working CTO — included at no extra cost, useful for reliability strategy, architecture reviews, or pressure-testing an incident process. Committing to 3+ seats or 12 months also earns 5–10% off.

SRE seats are on the standard rate card.

Same flat monthly rate, same includes — office, equipment, HR, management, replacement guarantee.