All case studies
HR-Tech

HR-Tech Company Case Study

99.99% uptime, <200ms job search latency
01

Amazon EKS Implementation

We implemented Amazon EKS to manage the majority of HR-Tech company's workloads and core data flows, handling over 80% of the platform's containerized operations in the ap-south-1 region. This included the primary job search and recruitment microservices which we designed as the backbone of user interactions, data processing, and integrations with AWS services like S3 for resume storage, DynamoDB for job profiles, Lambda for event-driven notifications, SQS for async workflows, SNS for pub/sub messaging, SES for transactional emails, RDS for reporting, CloudFront for CDN delivery, and Elastic Load Balancing for traffic distribution. We configured EKS's data plane with managed node groups using EC2 instances for persistent compute-intensive tasks like job matching algorithms and Fargate profiles for serverless bursty workloads like resume parsing, ensuring scalability across multi-AZ private subnets within the VPC.

02

Key Workloads Managed via EKS

We managed critical workloads including job search and candidate profile services, communication services handling email/SMS/WhatsApp notifications, event tracking systems, general queuing services, email processors, and webhook handlers. All deployments maintained multiple replicas for high availability with Horizontal Pod Autoscaler scaling based on CloudWatch CPU and memory metrics from Container Insights. Traffic flowed securely through CloudFront edge caching to Elastic Load Balancing which distributed requests across EKS services before reaching backend storage layers like DynamoDB, RDS, and S3.

03

Node Tagging Strategy & Compliance Alignment

We aligned HR-Tech company's AWS environment under our AWS Organization with a dedicated management account where Service Control Policies enforced mandatory resource tagging, denying EC2 launches without required tags. In HR-Tech company's member account, we tagged all EKS data plane nodes — EC2 instances in managed node groups — consistently with our enterprise tagging strategy. This supported cost allocation through Cost Explorer, compliance monitoring via AWS Config rules, and operational automation across all services.

04

Tag Definitions & Purpose

  • Environment tags set as "Production" for prod nodes or "Staging" for test environments, enabling lifecycle management like auto-termination of dev resources after 30 days.
  • Project tags identifying core recruitment platform infrastructure, facilitating Cost Explorer grouping.
  • Owner tags ensuring accountability for audits and incident response.
  • CostCenter tags aligning billing requirements with SCP policies.
  • Additional tags for workload type, compliance standard (GDPR), and auto-scaling group linkage.
05

Tagging Validation & Compliance

We verified EKS node tagging via the AWS CLI, with tags applied automatically through CloudFormation TagSpecifications in AWS::EKS::Nodegroup resources, meeting SCP requirements for non-empty string values on mandatory keys. We achieved 100% compliance validated quarterly through AWS Config rule ec2-instance-tagging-required, GuardDuty exception detection, CloudTrail audit trails, and CloudWatch Logs retention. WAF protected application endpoints from exploits while Secrets Manager handled automatic credential rotation and KMS ensured data encryption at rest across S3, RDS, and EKS etcd.

06

Deployment Verification — kubectl Outputs

We captured deployment outputs from the production EKS cluster running a recent Kubernetes version, confirming all replicas ready with HPA scaling policies active. Key deployments — including the BFF and candidate-backend services, AI-search and email consumers, the recruiter email processor, the event-tracker service, and the general queuing service — all reported healthy ready status, with the event-tracker service scaled to multiple replicas to absorb peak load.

07

Deployment Notes & Monitoring

All deployments are configured with multiple replicas for high availability, with the event-tracker deployment auto-scaling up to 10 replicas during peak loads via CloudWatch HPA. Pods used consistent labeling and versioning with defined CPU and memory resource requests. Fargate profiles handled serverless workloads like authentication services while EC2 nodes hosted stateful services with EBS volumes accessed via Route 53 private hosted zones. No deployment failures were detected, with monitoring showing less than 1 restart per hour across CloudWatch, GuardDuty, and WAF security layers.

08

Infrastructure Automation & Deployment Practices

We automated EKS infrastructure changes using Terraform for provisioning VPC configurations, EKS clusters, EC2 node groups, RDS instances, S3 buckets, Lambda functions, and IAM roles. CI/CD pipelines orchestrated application deployments where container images built and pushed to ECR automatically triggered updates to Kubernetes manifests and Helm charts applied to EKS. Rollback strategies included Terraform state reversion for infrastructure and Kubernetes Deployment rollbacks to previous ReplicaSets triggered by failed health checks. Production changes prohibited direct AWS Management Console access, with all operations enforced through pipelines monitored by CloudWatch alarms.

09

Key AWS Services Integration

For storage and database layers we deployed DynamoDB for sub-millisecond job profile queries, S3 with KMS encryption for resume and asset storage, and RDS with multi-AZ failover for enterprise reporting. Messaging infrastructure utilized SQS for decoupled async workflows, SNS for cross-service pub/sub notifications, and SES for high-volume transactional emails to recruiters and candidates. Security stack included GuardDuty for continuous threat detection across VPC Flow Logs and CloudTrail, WAF rulesets protecting web endpoints, CloudTrail capturing all API activity, Secrets Manager for credential rotation, and KMS for envelope encryption. Networking comprised VPC with multi-AZ private/public subnets, CloudFront global CDN caching, Route 53 DNS with health checks, and Elastic Load Balancing for intelligent traffic routing. Observability leveraged CloudWatch for metrics and logs with Container Insights, Athena for S3 log analysis, and Cost Explorer for spend optimization across all services.

10

Business Impact & Results

  • Job search latency under 200ms during peak concurrent user spikes.
  • 99.99% uptime through multi-AZ deployments and automated failover.
  • 40% infrastructure cost savings via EC2 Spot Instances, Fargate serverless scaling, and Cost Explorer optimizations.
  • Zero security breaches, with GuardDuty blocking over 1,200 threats monthly and full GDPR compliance maintained.
  • Deployment cycles reduced from weeks to under one hour.
  • Incident response times dropped from 30 minutes to 5 minutes via CloudWatch alerting and automated remediation.

Ready to achieve similar results?

Talk to our engineers about your cloud challenge. We'll get back to you within one business day.

Get in touch

Trusted by forward-thinking teams

Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo