Vanamala Venkataswamy
Ph.D. CS

Vanamala Venkataswamy

System Architect for the AI Era

Principal-level Systems Architect and Research Scientist specializing in building large-scale distributed computing, HPC, and data orchestration systems for AI/ML, simulation and research workloads. Specialized in cloud-native infrastructure that accelerates AI/ML pipelines.

Distributed Systems Cloud Computing Applied AI/ML DevOps/MLOps/Infrastructure

Technical Arsenal

Cloud & Infrastructure

Amazon AWS (EC2, S3, CloudFormation, SageMaker), Azure Cloud, Docker, Kubernetes

ML/AI Frameworks

PyTorch, TensorFlow, Scikit-learn, Matplotlib, Seaborn, Pandas

Languages

Python, Java, C, C++, Shell scripting

DevOps

CI/CD, Github, Ansible, Jenkins, Puppet, MLOPs, LLMOps

Certifications

Databricks Fundamentals, Databricks Advanced Machine Learning Operations, MLOps/LLMOps

Patents & Innovation

Power Aware Scheduling PDF

U.S. Patent No. 17/402.175

Granted Feb 2022

Professional Service & Outreach

Technical Committees

  • Technical Program Committee Member: IPDPS 2023/2024 (System Software Track)
  • Poster Session Judge: CAPWIC 2022
  • Technical Judge: Virginia State Science and Engineering Fair (VSSEF) 2021
  • Invited Panelist: "Machine Learning in Computing Systems," MLCS Workshop at SC’20

Mentorship & Community

  • Coach: First Lego League (FLL) for middle and elementary school students (2024–2025)
  • Mentor: Mentoring Ph.D. candidate, University of Virginia (2021–2022/2023-2026)
  • Representative: Society of Women Engineers (GradSWE) at UVA (2017)
  • Volunteer: Charlottesville Girls Geek Day & UVA Open House (2016)

Professional Experience

Technical Director & AI Advisor

2024—Present

AI/ML & Digital Product Development

  • Architecting resilient, full-stack website and app for real-estate solutions.
  • Establishing robust CI/CD pipelines and automated testing frameworks for the website and app.
  • Advising on Generative AI model fine-tuning, benchmarking, evaluation, and metric development.

Postdoctoral Researcher

2023—2024

Biocomplexity Institute, University of Virginia

  • Deployed a distributed orchestration platform, on-prem and cloud, for 10,000+ concurrent epidemiological simulations.
  • Benchmarked Generative AI surrogate models for High-Energy Physics simulation.

Intelligent Infrastructure Optimization

2016—2023

Ph.D. Research, University of Virginia

  • Developed RL-based schedulers for power-modulated datacenters, increasing profit by 6%-18%.
  • Built a modular simulator for datacenter resource management that supports heuristic and ML scheduling policies and renewable energy sources.

Distributed Systems Software Intern

May-2019 — Aug-2019

Lancium

  • Engineered a seamless scheduling pipeline by configuring the cloud backend and implementing custom software to intelligently match user job requests with available GPU resources.
  • Accelerated product release cycles by developing an automated testing and CI/CD pipeline, facilitating the beta release and maintenance of Lancium’s cloud platform.

Senior Applications Programmer Ananlyst

2011—2016

University of Virginia

  • National Grid Orchestration: Administered and optimized the Cross Campus Grid (XCG) and XSEDE grid infrastructure, architecting the integration of compute and storage resources across multiple national supercomputing facilities.
  • Systems Integration & Tooling: Engineered the GenesisII software ecosystem by developing grid command-line tools and web services, while establishing a comprehensive unit and regression testing framework to ensure platform stability.
  • Technical Leadership & Incident Response: Served as the primary technical lead and point of contact for nationwide grid users, directing incident response and troubleshooting efforts for complex production issues within high-performance computing environments.

Member Technical Staff 2

2005—2008

Center for Development of Advanced Computing

  • Ported 32-bit Message Passing Interface (MPI) implementation to 64-bit architecture, implementing and testing MPI for TCP/IP and Virtual Interface Architecture (VIA) communication protocols across distributed systems.
  • Designed and implemented Parallel File System using NFS3 protocol and MPICH, supporting AIX, Linux, and SunOS platforms with multithreaded IO servers for high-performance parallel processing.
  • Built grid computing infrastructure for Garuda Grid project, implementing Storage Resource Broker (SRB) and iRODS testbed for distributed data management across compute clusters.
  • Integrated Cluster Management System (CMS) with LoadLeveller for distributed job scheduling and workload management.

Selected Projects & Research

CaloBench Research

CaloBench: Generative Models for Calorimeters

A comprehensive benchmark study evaluating Generative AI surrogate models (GANs, VAEs, Diffusion) against Geant4 simulation samples for high-energy physics. Focuses on data correlation and performance metrics for the ATLAS Calorimeter.

RL Scheduling

Launchpad & RARE: RL-Based Scheduling

Developed Reinforcement Learning agents for intelligent resource management in datacenters. Using offline and online RL to optimize for power awareness and renewable energy utilization, achieving up to 18% profit increase.

AWS cloud orchestration

Parallel Compute Engine (PaCE)

Implemented a distributed parallel computing system designed to execute computationally intensive tasks across multiple nodes. It utilizes a master-worker architecture to manage task distribution, synchronization, and resource allocation to improve processing efficiency.

Selected Scholarly Contributions

Bench 2025 PDF

CaloBench: A Benchmark Study of Generative Models for Calorimeter Showers

JSSPP 2024 PDF

Launchpad: Learning to Schedule Using Offline and Online RL Methods

Ph.D. Dissertation, UVA 2023 PDF

Scheduling to Ensure Performance and Cost Effectiveness in Power-Modulated Datacenters Project Page

JSSPP 2022 PDF

RARE: Renewable Energy Aware Resource Management in Datacenters

MLCS 2020 PDF

Scheduling in Data Centers Running on Renewable Energy with Deep Reinforcement Learning