Omnibond cloud HPC: Clemson’s 1.1M vCPU Spot Fleet for large-scale topic modeling.

Natural language processing (NLP) is revolutionizing how we uncover insights from vast text corpora, from predicting business trends to informing public policy. But training topic models at scale demands computational firepower beyond most on-premises setups. Clemson University's School of Computing, led by Professor Amy Apon and her team, shattered records by launching the largest cloud-based high-performance cluster in a single AWS region: 1,119,196 vCPUs across thousands of EC2 Spot Instances. This wasn't just a flex, it enabled nearly half a million parallel experiments on topic modeling, analyzing 17 years of computer science journal abstracts (533,560 documents, 32.5M words) and NIPS Conference papers (2,484 documents, 3.3M words). Outputs streamed to Amazon S3 for deep analysis on model convergence, topic quality, and parameter impacts.
The result? A peer-reviewed study that optimized topic modeling for real-world applications, proving AWS could extend Clemson's on-prem Palmetto Cluster without disruption. At the heart of this achievement was Omnibond®'s hybrid orchestration technology, turning complex Spot Fleet management into seamless, autoscaling workflows. Here's how we made it happen.
Clemson's Palmetto supercomputer handles routine research, but massive parameter sweeps for topic models, like testing hyperparameters across datasets, require burst capacity that's impossible on-campus. The team needed to:
Manual cloud setups meant weeks of scripting, risk of Spot interruptions, and no hybrid safety net. Clemson required a toolset that automated provisioning, handled preemptions, and scaled elastically, without YAML nightmares or vendor silos.
Omnibond® partnered with Clemson and AWS to deliver a turnkey framework, leveraging our deep HPC expertise to bridge on-prem and cloud. Key components included:
Omnibond®'s hands-on collaboration was pivotal: We optimized PAW for AWS, providing expertise in Spot management and hybrid integration that let Clemson focus on science, not plumbing.
Launched in US East (N. Virginia) in a single region, the cluster peaked at 1,119,196 vCPUs, rivaling the world's top supercomputers. Highlights:
Professor Apon raved: "I am absolutely thrilled with the outcome... They used resources from AWS and Omnibond® and developed a new software infrastructure to perform research at a scale and time-to-completion not possible with only campus resources. Per-second billing was a key enabler."
The project slashed costs to a fraction of on-demand pricing, freeing NSF-funded resources for innovation. It yielded breakthroughs in topic modeling, quantifying how parameters influence convergence and quality, applicable to AI forecasting, policy analysis, and more. Published in peer-reviewed journals, it showcased AWS + Omnibond® as a blueprint for hybrid HPC.
This run is an Omnibond® company capability. The scale experience, hybrid workflows, and data-adjacent AI work inform projectEureka: project-centric browser workspaces for AI and simulation work on AWS, GCP, or on-prem Kubernetes.
Sources: Based on AWS's blog on Clemson's NLP project, highlighting Omnibond®'s technology and cloud integration.
Explore research and education or enterprise paths for projectEureka™ workspaces.