All case studies
FinTech / AI

Leading Fintech Company

50% Reduction in GPU Costs
01

Scalable Multi-GPU On-Prem LLM Deployment for Fintech

CLIENT OVERVIEW: A leading fintech company specializing in AI-driven risk assessment, fraud detection, and automated financial advisory services needed a robust on-prem infrastructure to support its expanding AI capabilities. Their models process vast amounts of transactional data, identifying anomalies and providing real-time insights to financial institutions. To enhance performance and scalability, they required an efficient way to deploy multiple open-source large language models (LLMs) on NVIDIA GPUs. The primary goal was to ensure seamless auto-scaling and optimized inference performance for real-time financial applications.

02

Challenges Faced

  • 1) High Cloud Costs for Sustained Workloads: While cloud-based GPU instances provided initial flexibility, the long-term costs of running LLMs at scale became unsustainable. The unpredictable nature of cloud expenses made it difficult to optimize budgets, prompting the need for a cost-effective, on-prem solution.
  • 2) Ensuring Data Security & Privacy: With AI models processing highly confidential financial transactions, data security was a top priority. Relying on cloud providers introduced risks related to data exposure and third-party access.
  • 3) Meeting Regulatory Compliance Requirements: The fintech sector is subject to stringent compliance standards such as GDPR, PCI-DSS, and financial data residency regulations. Storing and processing data in the cloud raised concerns about regulatory adherence.
  • 4) Optimizing Performance for Real-Time Inference: AI-driven fraud detection and financial advisory services demand low-latency inference and seamless scalability. Cloud-based environments introduced performance bottlenecks and networking constraints.
03

Infrastructure Setup on OpenShift

  • To overcome the challenges of cost efficiency, security, compliance, and performance, we designed and implemented a production-grade OpenShift Kubernetes cluster with a multi-GPU architecture.
  • Bare-metal Deployment: We installed OpenShift Container Platform on high-performance bare-metal servers equipped with NVIDIA A100 GPUs. This eliminated recurring cloud costs and provided full control over infrastructure.
  • NVIDIA GPU Operator: We integrated the NVIDIA GPU Operator for seamless driver installation, GPU monitoring, and optimized scheduling within Kubernetes, ensuring high efficiency.
  • Storage Configuration: We used CephFS for scalable, high-availability storage, allowing persistent storage for large model weights and checkpoints.
  • Service Mesh Implementation: Istio service mesh was configured to enable secure traffic management, encryption, and observability across AI workloads.
  • High-Speed Networking: We implemented RDMA over Converged Ethernet (RoCE) to reduce latency between GPU nodes.
04

GPU Optimization & Deployment of LLM Models

  • CUDA Toolkit & cuDNN: We installed NVIDIA’s CUDA toolkit, including cuDNN and TensorRT, to maximize deep learning inference performance while reducing resource consumption.
  • MIG (Multi-Instance GPU): We enabled NVIDIA MIG to partition A100 GPUs, allowing multiple AI models to run efficiently on a single GPU, reducing hardware costs.
  • GPUDirect RDMA: We implemented GPUDirect RDMA to bypass CPU bottlenecks, ensuring ultra-low latency and high-bandwidth GPU communication.
  • NVIDIA Triton Inference Server: We deployed NVIDIA Triton, enabling dynamic batching and model optimization to serve multiple LLM models with high throughput.
  • Deployed Models: LLaMA 2 (optimized for financial document analysis), Mistral-7B (for conversational AI chatbots), Bloom-176B (for compliance automation), and Dolly-2.0 (for NLP tasks).
05

Auto-Scaling Mechanism & Monitoring

  • Horizontal Pod Autoscaler (HPA): We configured Kubernetes HPA to dynamically scale inference pods based on GPU utilization.
  • KEDA Integration: We integrated Kubernetes Event-Driven Autoscaler (KEDA) to trigger scaling based on workload spikes, preventing downtime during peak transaction periods.
  • Cluster Autoscaler: OpenShift Cluster Autoscaler was deployed to automatically add or remove GPU nodes as demand fluctuated.
  • Zero Trust Security Model: We enforced strict RBAC policies with identity-based access control.
  • Threat Detection & Log Aggregation: Deployed Falco for runtime security monitoring. Used Elasticsearch, Fluentd, Prometheus, Grafana, Lowkey and Kibana (EFK) to centralize logging.
  • Fault Tolerance & Disaster Recovery: Multi-master OpenShift cluster, redundant GPU nodes with automatic failover, and Velero for periodic backups.
06

Results

  • 5x Faster Inference Speeds: By leveraging CUDA optimizations, model quantization, and parallel execution, we significantly accelerated inference speeds.
  • 50% Reduction in GPU Costs: Optimized GPU utilization using NVIDIA MIG and dynamic auto-scaling, cutting GPU-related expenses in half.
  • Seamless, On-Demand Scaling: Enabled the infrastructure to scale dynamically from 5 to 100+ GPUs in response to fluctuating workloads.
  • 99.99% Uptime & High Availability: Implemented a fault-tolerant OpenShift architecture with redundant GPU nodes and disaster recovery mechanisms.
  • Enterprise-Grade Security & Compliance: Achieved strict compliance with financial and AI security standards through zero-trust architecture, real-time threat monitoring (Falco), and automated compliance enforcement (OpenSCAP).
07

Conclusion

We enabled Fintech to deploy, scale, and optimize multiple LLMs in a secure, high-performance on-premises environment. With OpenShift, NVIDIA CUDA toolkit, and advanced auto-scaling, Fintech achieved unparalleled GPU efficiency, ensuring their AI workloads were production-ready for enterprise-scale deployments.

Ready to achieve similar results?

Talk to our engineers about your cloud challenge. We'll get back to you within one business day.

Get in touch

Trusted by forward-thinking teams

Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo
Client Logo