Leading Fintech Company
Scalable Multi-GPU On-Prem LLM Deployment for Fintech
CLIENT OVERVIEW: A leading fintech company specializing in AI-driven risk assessment, fraud detection, and automated financial advisory services needed a robust on-prem infrastructure to support its expanding AI capabilities. Their models process vast amounts of transactional data, identifying anomalies and providing real-time insights to financial institutions. To enhance performance and scalability, they required an efficient way to deploy multiple open-source large language models (LLMs) on NVIDIA GPUs. The primary goal was to ensure seamless auto-scaling and optimized inference performance for real-time financial applications.
Challenges Faced
- 1) High Cloud Costs for Sustained Workloads: While cloud-based GPU instances provided initial flexibility, the long-term costs of running LLMs at scale became unsustainable. The unpredictable nature of cloud expenses made it difficult to optimize budgets, prompting the need for a cost-effective, on-prem solution.
- 2) Ensuring Data Security & Privacy: With AI models processing highly confidential financial transactions, data security was a top priority. Relying on cloud providers introduced risks related to data exposure and third-party access.
- 3) Meeting Regulatory Compliance Requirements: The fintech sector is subject to stringent compliance standards such as GDPR, PCI-DSS, and financial data residency regulations. Storing and processing data in the cloud raised concerns about regulatory adherence.
- 4) Optimizing Performance for Real-Time Inference: AI-driven fraud detection and financial advisory services demand low-latency inference and seamless scalability. Cloud-based environments introduced performance bottlenecks and networking constraints.
Infrastructure Setup on OpenShift
- To overcome the challenges of cost efficiency, security, compliance, and performance, we designed and implemented a production-grade OpenShift Kubernetes cluster with a multi-GPU architecture.
- Bare-metal Deployment: We installed OpenShift Container Platform on high-performance bare-metal servers equipped with NVIDIA A100 GPUs. This eliminated recurring cloud costs and provided full control over infrastructure.
- NVIDIA GPU Operator: We integrated the NVIDIA GPU Operator for seamless driver installation, GPU monitoring, and optimized scheduling within Kubernetes, ensuring high efficiency.
- Storage Configuration: We used CephFS for scalable, high-availability storage, allowing persistent storage for large model weights and checkpoints.
- Service Mesh Implementation: Istio service mesh was configured to enable secure traffic management, encryption, and observability across AI workloads.
- High-Speed Networking: We implemented RDMA over Converged Ethernet (RoCE) to reduce latency between GPU nodes.
GPU Optimization & Deployment of LLM Models
- CUDA Toolkit & cuDNN: We installed NVIDIA’s CUDA toolkit, including cuDNN and TensorRT, to maximize deep learning inference performance while reducing resource consumption.
- MIG (Multi-Instance GPU): We enabled NVIDIA MIG to partition A100 GPUs, allowing multiple AI models to run efficiently on a single GPU, reducing hardware costs.
- GPUDirect RDMA: We implemented GPUDirect RDMA to bypass CPU bottlenecks, ensuring ultra-low latency and high-bandwidth GPU communication.
- NVIDIA Triton Inference Server: We deployed NVIDIA Triton, enabling dynamic batching and model optimization to serve multiple LLM models with high throughput.
- Deployed Models: LLaMA 2 (optimized for financial document analysis), Mistral-7B (for conversational AI chatbots), Bloom-176B (for compliance automation), and Dolly-2.0 (for NLP tasks).
Auto-Scaling Mechanism & Monitoring
- Horizontal Pod Autoscaler (HPA): We configured Kubernetes HPA to dynamically scale inference pods based on GPU utilization.
- KEDA Integration: We integrated Kubernetes Event-Driven Autoscaler (KEDA) to trigger scaling based on workload spikes, preventing downtime during peak transaction periods.
- Cluster Autoscaler: OpenShift Cluster Autoscaler was deployed to automatically add or remove GPU nodes as demand fluctuated.
- Zero Trust Security Model: We enforced strict RBAC policies with identity-based access control.
- Threat Detection & Log Aggregation: Deployed Falco for runtime security monitoring. Used Elasticsearch, Fluentd, Prometheus, Grafana, Lowkey and Kibana (EFK) to centralize logging.
- Fault Tolerance & Disaster Recovery: Multi-master OpenShift cluster, redundant GPU nodes with automatic failover, and Velero for periodic backups.
Results
- 5x Faster Inference Speeds: By leveraging CUDA optimizations, model quantization, and parallel execution, we significantly accelerated inference speeds.
- 50% Reduction in GPU Costs: Optimized GPU utilization using NVIDIA MIG and dynamic auto-scaling, cutting GPU-related expenses in half.
- Seamless, On-Demand Scaling: Enabled the infrastructure to scale dynamically from 5 to 100+ GPUs in response to fluctuating workloads.
- 99.99% Uptime & High Availability: Implemented a fault-tolerant OpenShift architecture with redundant GPU nodes and disaster recovery mechanisms.
- Enterprise-Grade Security & Compliance: Achieved strict compliance with financial and AI security standards through zero-trust architecture, real-time threat monitoring (Falco), and automated compliance enforcement (OpenSCAP).
Conclusion
We enabled Fintech to deploy, scale, and optimize multiple LLMs in a secure, high-performance on-premises environment. With OpenShift, NVIDIA CUDA toolkit, and advanced auto-scaling, Fintech achieved unparalleled GPU efficiency, ensuring their AI workloads were production-ready for enterprise-scale deployments.
Ready to achieve similar results?
Talk to our engineers about your cloud challenge. We'll get back to you within one business day.
Trusted by forward-thinking teams















