preeloader
Qubow Logo
Full time job

AWS Infrastructure & DevOps Lead

Job Summary

1. Role Purpose

Qubow is seeking a hands-on AWS Infrastructure & DevOps Lead to take end-to-end operational ownership of our cloud infrastructure across Production, UAT, Development, Testing, and Staging environments. This role directly covers the full scope of responsibilities that would otherwise be outsourced to a managed services vendor — from provisioning and deployment to monitoring, incident response, security operations, and reporting.

The ideal candidate is a hybrid of AWS Solutions Architect and DevOps Engineer who can operate independently, build standardized CI/CD practices, and give Qubow's leadership visibility and control over infrastructure without external dependency.

Key Responsibilities

A. New Project Deployment & Provisioning

  • Set up new AWS accounts, VPCs, subnets, security groups, route tables, NAT/Internet gateways
  • Manage IAM roles, policies, users, and key management following least-privilege principles
  • Provision and configure compute services: EC2, ECS, EKS, Lambda, Auto Scaling Groups
  • Provision databases: Aurora, MySQL, PostgreSQL, RDS, Redis, ElastiCache
  • Configure DNS (Route53), SSL certificates, load balancers, and CloudFront CDN
  • Own Infrastructure-as-Code practices using Terraform, CloudFormation, or CDK — all infrastructure version-controlled

B. CI/CD & Environment Management

  • Build and maintain CI/CD pipelines using GitHub Actions, CodePipeline, CodeBuild, CodeDeploy (Jenkins where applicable)
  • Support deployments across Dev, QA, UAT, and Production with defined approval workflows
  • Maintain consistency across environments

C. Daily Server & Application Administration

  • Linux administration: patching, security updates, package upgrades
  • Server lifecycle management: creation, resizing, storage expansion, AMI management, instance optimization
  • Application support: deployment assistance, service restarts, log analysis, troubleshooting, configuration changes
  • SSL renewal and certificate installation
  • User/access management: IAM, SSH keys, VPN access, role and permission management

D. Deployment Management

  • Manage branch deployments, hotfixes, rollbacks, and release coordination
  • Implement blue-green and zero-downtime deployment strategies
  • Perform deployment validation before and after releases

E. Monitoring & Observability

  • Implement and maintain monitoring via CloudWatch, Grafana, Prometheus, ELK/OpenSearch, and AWS Health Dashboard
  • Track CPU, memory, disk, network, database, application health, API failures, SSL/certificate expiry, storage, backups, replication, and latency

F. Incident Management

  • Own detection, triage, root cause analysis, resolution, and escalation of incidents
  • Maintain a priority matrix (P1–P4) and produce Post-Incident Reports
  • Meet defined response/resolution targets (see SLA table below)

G. Backup & DR Operations

  • Verify and monitor backups, run restore tests, manage snapshots and retention policies
  • Resolve backup failures promptly

H. Security Operations

  • Manage security patching, IAM reviews, least-privilege enforcement, key rotation, certificate management
  • Review security groups; remediate vulnerabilities
  • Act on AWS Trusted Advisor, Security Hub, and GuardDuty findings; manage WAF rules

I. Performance & Cost Optimization

  • Continuously right-size EC2/RDS, optimize storage, caching, CloudFront, and Auto Scaling
  • Recommend Reserved Instances and cost-saving opportunities

J. Network Appliance Setup & Management

  • Set up, configure, and manage inline network appliances (bridge/relay-based devices sitting between router and switch) used for network intelligence, traffic control, and WiFi monetisation
  • Configure and maintain traffic shaping rules and passive network monitoring on these appliances
  • Manage firmware/software update workflows and cloud-to-device sync/backup flows for deployed edge devices

K. Documentation & Reporting

  • Maintain architecture diagrams, deployment guides, recovery procedures, runbooks, and SOPs
  • Produce weekly, monthly, and quarterly reports covering incidents, deployments, capacity, security, cost optimization, backups, and SLA compliance
3. What Stays with Leadership (Not This Role)

To keep scope clear, the following remain with the CTO/Founders rather than this position:

  • Infrastructure architecture and technology roadmap decisions
  • Cloud strategy and platform standards
  • Security policy definition
  • Disaster recovery strategy (this role executes DR operations, not strategy)
  • R&D and new technology evaluation
  • Final approval of production architecture changes
4. Service Level Expectations
Severity Response Time Resolution Target
Critical 15 minutes 2 hours
High 30 minutes 4 hours
Medium 2 hours 1 business day
Low 4 hours 3 business days

Target infrastructure availability: 99.9% uptime

5. Required Qualifications & Experience
  • Minimum 5-8 years of hands-on AWS infrastructure/ DevOps experience, including production workload support
  • AWS Certified Solutions Architect and/ or AWS DevOps Engineer certification (strongly preferred)
  • Strong Linux administration background (see detailed Linux proficiency requirements below)
  • Proven experience owning CI/CD pipelines end-to-end
  • Experience with 24×7 or on-call support models
  • Prior experience acting as a technical lead or single point of accountability for infrastructure (agency, vendor, or in-house)
6. Linux Proficiency Requirements

Beyond standard server administration, the candidate must be able to independently set up and manage inline network appliances (edge devices running embedded Linux, positioned between a router and switch for network intelligence, traffic control, and bandwidth monetisation). Required depth:

  • Core administration: Debian/Ubuntu-based systems administration — package management, systemd service management, kernel modules, boot process, log management (journald/syslog)
  • Networking internals: Deep comfort with Linux networking — bridging (brctl/bridge utilities), VLANs, routing tables, network namespaces, and inline/transparent bridge or pass-through/bypass relay configurations
  • Firewall & traffic control: iptables/nftables for security and traffic interception; tc (traffic control) and queueing disciplines for bandwidth shaping and tiered monetisation
  • Network intrusion/monitoring tooling: Working knowledge of Suricata (rules, eve.json output) and log shipping via Filebeat into ELK/ OpenSearch for passive network intelligence and compliance capture
  • Embedded/appliance Linux: Experience with resource-constrained/ embedded Linux builds, remote firmware/ software updates without service interruption, and local-to-cloud data sync patterns (local DB syncing back to a central service)
  • Router/ AP integration: Familiarity with MikroTik RouterOS (API/ CLI) or equivalent router API integration, to support both inline enforcement and direct router-API-based control modes
  • Scripting: Strong Bash and Python for device provisioning scripts, health-check agents, and automation glue between edge appliances and the cloud backend
  • Captive portal / RADIUS familiarity: Exposure to captive portal flows, RADIUS/AAA, and OTP-gated guest access flows is a strong plus, for venue-based guest WiFi management
7. Preferred Technical Skills

AWS · GitHub & GitHub Actions · Terraform · CloudFormation · Docker · ECS · Linux (incl. embedded/ edge) · Bash · Python · Ansible · Prometheus · Grafana · ELK/ OpenSearch · CloudWatch · Suricata · iptables/nftables · MikroTik RouterOS · RADIUS/ Captive Portal

8. Soft Skills/ Working Style
  • Comfortable operating with a high degree of autonomy across multiple environments and projects (Qubow runs several concurrent ventures/ products)
  • Strong documentation discipline — this role is expected to leave clear runbooks and audit trails
  • Clear, proactive communicator; comfortable presenting monthly operational and security reports to leadership
  • Able to triage and prioritize across competing incidents/requests without constant supervision
9. Success Criteria (First 6–12 Months)
  • All new project deployments run on standardized, documented CI/CD pipelines
  • Routine infrastructure requests handled independently with minimal escalation
  • Monitoring/alerting shifts from reactive to proactive; incidents resolved within SLA
  • Infrastructure availability consistently at or above 99.9% uptime
  • Monthly reports delivered on schedule, covering operations, incidents, cost, and security posture
  • All infrastructure changes documented, reviewed, and version-controlled
  • Leadership able to redirect time from operational firefighting to architecture, governance, and R&D
10. Compensation & Contract
  • To be defined based on candidate experience level and whether the role is structured as full-time employment or a retained contract
  • Performance reviewed at 1-month probation checkpoint, then quarterly