Executive Summary
Building reliable applications on distributed, interruptible infrastructure is a solved engineering problem. Companies like Civitai (10M images/day), Undetectable AI (tens of millions of users), Blend (6,000 images/hour), Klyne AI (thousands of molecular simulations), and Theres-An-AI-For-That (TAAFT) (millions of users) have successfully deployed production workloads on SaladCloud while achieving 50-85% cost savings compared to traditional clouds.
Key Benefits:
-
Cost Reduction: 50-85% savings vs hyperscalers
-
Infinite Scale: Access to a huge supply of nodes globally with transparent and predictable pricing
-
Production Ready: Powers some of the internet's largest AI applications
It's no secret that SaladCloud is comprised of thousands of privately owned and interruptible PCs and servers around the world, and if you're newer to building at scale, you may find yourself wondering, "How on earth could you possible build a reliable application on top of that?"
While this challenge does present itself at smaller operating scales on SaladCloud than on hyperscale clouds like AWS or Azure, it is not a unique problem to building on SaladCloud, or even on distributed clouds generally. This is good news because it means there are a number of well-established, tried-and-true solutions that have been developed over decades in the industry at large.
The Technical Challenge
Before we dive into solutions, let's examine the problem in more detail.
Any individual container instance can go down without warning. On SaladCloud this tends to be either the workload getting preempted by a higher priority workload (SaladCloud uses priority pricing with four tiers from Batch to High), or the Chef (compute host) reclaiming their GPU for another purpose.
This same challenge exists across all distributed systems: On AWS or Azure, it may be due to failures in the underlying hardware, service outages, or preemption by AWS Spot instances and Azure Spot VMs. Google Cloud's preemptible instances can be terminated with 30 seconds notice, similar to how SaladCloud instances can be reclaimed. Even Kubernetes clusters running on traditional clouds must handle pod evictions, node failures, and rolling updates that terminate containers without warning.
The fundamental reality is that any distributed system must assume individual compute units are unreliable. This principle has shaped decades of distributed systems engineering, from the early days of cluster computing to modern microservices architectures.
Despite this, end users expect a service to be largely uninterrupted, but not completely uninterrupted. Site reliability engineering tells us it is not possible to build a system with 100% uptime. Instead, the goal is to balance user experience, cost, and real-world constraints. These goals are expressed through Service Level Objects and Agreements (SLOs, SLAs), typically in the form of availability percentages like 99.9% uptime.
The Solution: Redundancy and Graceful Degradation
The fundamental approach to achieving high uptime on any distributed system, whether its AWS or SaladCloud, is redundancy. Instead of relying on a single instance to handle your workload, you deploy multiple instances across different hosts and regions.
These patterns are foundational to distributed systems engineering:
-
Netflix pioneered "Chaos Engineering" and tools like Chaos Monkey specifically because they run on AWS instances that can fail at any time
-
Google's internal infrastructure assumes thousands of server failures daily across their global fleet
-
Kubernetes was designed from the ground up to handle pod failures and node outages gracefully
-
Apache Kafka, Cassandra, and other distributed databases use replication and consensus algorithms to maintain availability despite individual node failures
1. Horizontal Scaling with Load Balancing

Deploy multiple replicas of your application across different SaladCloud nodes. Use a load balancer to distribute incoming requests across these instances. When one instance goes down, the load balancer automatically routes traffic to healthy instances.
This is the same pattern used everywhere in distributed systems:
-
AWS Auto Scaling Groups deploy instances across multiple availability zones with health checks
-
Kubernetes ReplicaSets maintain desired pod counts and replace failed instances automatically
-
Content Delivery Networks (CDNs) like Cloudflare route traffic away from failed edge nodes
-
Database clusters use primary-replica configurations with automatic failover
Key considerations:
-
Leverage SaladCloud's truly global distribution - with nodes spanning continents, you can achieve geographic redundancy that many traditional clouds can't match at this scale
-
Deploy instances across multiple countries to minimize the impact of local outages
-
Use health checks to quickly detect and remove failed instances from the load balancer pool
-
Implement circuit breakers to prevent cascading failures
2. Stateless Application Design
Similar to building web services on any other cloud, design your applications for SaladCloud to be stateless whenever possible. This means storing session data, user state, and persistent data in external services like databases or caches rather than in memory on the compute instances.
Stateless design is a cornerstone of scalable architecture everywhere:
-
REST APIs are designed to be stateless, allowing any server to handle any request
-
Microservices architectures separate state management from business logic
-
Container orchestration platforms like Kubernetes, Docker Swarm, and ECS assume containers are stateless and ephemeral
-
Lambda functions and serverless computing are inherently stateless
Benefits:
-
Failed instances can be replaced without data loss
-
New instances can immediately handle any request without warm-up
-
Simplified scaling and deployment processes
3. Graceful Shutdown Handling
Implement proper shutdown procedures in your applications to handle preemption gracefully. This includes:
-
Listening for termination signals (SIGTERM)
-
Completing in-flight requests before shutting down
-
Saving any critical state to persistent storage
-
Implementing a grace period for cleanup operations
Graceful shutdown is a universal requirement:
-
Kubernetes pods receive SIGTERM signals during rolling updates and node drains
-
AWS Application Load Balancers stop sending new requests to instances during termination
-
Docker containers use SIGTERM for graceful shutdown in production environments
-
Process managers like systemd, PM2, and supervisord all implement graceful shutdown patterns
4. Queue-Based Processing
For batch processing or long-running tasks, implement a queue-based architecture. SaladCloud provides both native job queues and supports external queue systems like AWS SQS for tasks that may exceed the 100-second timeout limit of the container gateway. This approach provides several advantages:
-
Tasks can be retried if an instance fails mid-processing
-
Work can be redistributed to healthy instances
-
Progress isn't lost when instances are preempted
-
For long-running tasks, you can implement checkpointing to save progress to cloud storage and resume from the last checkpoint after interruptions
Queue-based processing is fundamental to distributed systems:
-
Message queues like RabbitMQ, Apache Kafka, and AWS SQS provide at-least-once delivery guarantees
-
Task queues like Celery, Sidekiq, and Google Cloud Tasks handle job distribution across worker processes
-
Stream processing systems like Apache Flink and Apache Storm checkpoint state to handle worker failures
-
Batch processing frameworks like Apache Spark and MapReduce automatically retry failed tasks on different nodes
Implementation pattern:

For detailed guidance on implementing this pattern, see the SaladCloud long-running tasks guide.
5. Monitoring and Auto-Scaling
Implement comprehensive monitoring to track instance health and automatically scale your deployment based on demand and availability:
-
Monitor instance health and response times
-
Set up alerts for when available capacity drops below thresholds
-
Implement auto-scaling rules to maintain desired redundancy levels
-
Use metrics to optimize your bidding strategy
Auto-scaling and monitoring are industry standards:
-
AWS Auto Scaling adjusts EC2 instance counts based on CloudWatch metrics
-
Google Cloud's Compute Engine provides managed instance groups with auto-scaling
-
Kubernetes Horizontal Pod Autoscaler scales pods based on CPU, memory, or custom metrics
-
Modern observability platforms like Datadog, New Relic, and Prometheus provide comprehensive monitoring for distributed systems
6. Data Persistence and Backup Strategies
Since SaladCloud instances are ephemeral, implement robust data persistence strategies:
-
Use external databases or object storage for persistent data
-
Implement regular backups of critical data
-
Consider using distributed storage systems for high availability
-
Cache frequently accessed data in memory stores like Redis
External storage is standard practice across all cloud architectures:
-
Twelve-Factor App methodology explicitly separates stateless processes from stateful backing services
-
Amazon RDS, Google Cloud SQL, and Azure Database provide managed database services designed for ephemeral compute
-
Object storage services like S3, Google Cloud Storage, and Azure Blob Storage are designed for durability across instance failures
-
Distributed caching systems like Redis Cluster and Memcached provide high availability caching
Implementation Decision Framework
No matter where you build—whether on SaladCloud, AWS, Google Cloud, or your own Kubernetes cluster—the fundamental decision between real-time and batch processing remains the same. The choice depends on your workload characteristics, not your infrastructure provider. These patterns have been refined across decades of distributed systems engineering and apply universally.
Choose Real-Time Inference (Container Gateway) when:
-
Response times under 100 seconds
-
Immediate results required (like Civitai's image generation)
-
Can handle request failures gracefully
-
Traffic patterns are predictable
Choose Queue-Based Processing when:
-
Tasks may exceed 100 seconds (like Klyne's molecular simulations)
-
High throughput more important than low latency
-
Complex workloads requiring checkpointing
-
Need guaranteed processing with automatic retries
Hybrid Approach (like TAAFT):
-
Use container gateway for lightweight text tools
-
Use job queues for compute-intensive image generation
-
Leverage different GPU types for different workload characteristics
This decision framework mirrors patterns used across the industry:
-
Netflix uses real-time services for video streaming and batch processing for recommendation algorithms
-
Uber combines real-time ride matching with batch processing for pricing optimization
-
Spotify uses real-time APIs for music playback and queue-based processing for playlist generation
Monitoring and Observability
Achieving high uptime requires visibility into your system's health:
-
Application metrics: Response times, error rates, throughput
-
Infrastructure metrics: CPU, memory, GPU utilization across instances
-
Business metrics: Active user sessions, revenue-impacting operations
-
Synthetic monitoring: Automated tests that verify functionality from user perspective
These observability practices are universal:
-
Google's SRE practices emphasize the four golden signals: latency, traffic, errors, and saturation
-
The Observability Engineering book by Honeycomb outlines these same principles for any distributed system
-
Prometheus and Grafana provide open-source monitoring stacks used across industries
-
OpenTelemetry standardizes observability across programming languages and platforms
Cost Optimization with Proven Results
The companies above demonstrate that building redundant systems doesn't have to break the bank:
-
Right-sizing for workload: Civitai uses consumer GPUs that deliver 4-8X more images per dollar than AI-focused datacenter GPUs for Stable Diffusion inference
-
Intelligent scaling: Klyne AI implements auto-scaling that monitors job queues and adjusts replica counts via SaladCloud's API, scaling from tens to thousands of GPUs based on simulation demands
-
Priority-based cost control: Use SaladCloud's 4-tier priority system to balance cost and availability - run batch workloads on lower-cost tiers while maintaining higher priorities for critical real-time services
Conclusion
Building reliable applications on SaladCloud requires embracing the distributed, ephemeral nature of the platform rather than fighting against it. The SaladCloud architectural overview shows how containers are orchestrated across thousands of consumer GPUs worldwide.
The engineering principles required are identical to those used by every major technology company:
-
Amazon built their entire retail platform assuming individual servers will fail
-
Google designed their search infrastructure to handle datacenter outages gracefully
-
Facebook (Meta) scales their social platform across millions of servers using these same redundancy patterns
-
Netflix famously practices "Chaos Engineering" to ensure their systems work despite constant failures
As demonstrated by Civitai's 10M daily images, Blend's 85% cost savings, and Klyne's burst scaling capabilities, implementing redundancy and designing for failure enables you to achieve high uptime that meets or exceeds traditional cloud offerings—often at a fraction of the cost.
The key is to think of individual instances as cattle, not pets. They're replaceable resources that will come and go, but your application as a whole should remain resilient and performant. With proper planning and implementation, SaladCloud's distributed GPU infrastructure can provide the reliability your users expect while delivering significant cost savings.
These patterns have been proven at scale across the industry. The same architectural principles that power Google's search, Netflix's streaming, and Amazon's e-commerce also enable reliable applications on SaladCloud. The tools and technologies may differ, but the fundamental approach to building resilient distributed systems remains consistent.
Getting Started Checklist:
-
Evaluate your workload characteristics - response time requirements, failure tolerance, scaling patterns
-
Start with a pilot - Deploy 3-5 replicas to test reliability patterns and performance
-
Implement monitoring - Track instance health, response times, and business metrics
-
Design for failure - Build retry logic, health checks, and graceful degradation
-
Scale gradually - Use proven patterns from successful deployments like those above
Remember, the goal isn't perfect uptime—it's the right balance of reliability, performance, and cost for your specific use case. Start with these patterns, monitor your results, and iterate based on real-world performance data.
Interested in SaladCloud's new tier with enterprise GPUs (H100s, A100s, L40S, and more)? Check out SaladCloud Secure.