Senior Linux Administrator — Infrastructure Engineering & Automation
Developed and deployed solutions in a large-scale, segmented Linux environment with restricted access between network segments. Configured the network connectivity and service access needed for automation, internal applications, and platform solutions.
- Designed and launched an AWS GPU/ML platform from scratch for approximately 15 ML engineers across two projects, supporting model training and inference, data storage, and ML pipelines. Used GPU-enabled EC2 instances, Aurora PostgreSQL, S3/EFS, and a containerized stack. Configured access controls, monitoring, and backups.
- Developed and deployed a Go client on over 10,000 Linux hosts to apply domain access rules from a centralized access management system.
- Designed and deployed production HAProxy infrastructure for critical payment services, including VTS and 3D Secure, with SSL termination and backup routing. This backup infrastructure later became the primary setup.
- Developed an internal Java inventory and tracking application integrated with Foreman, Ansible, CMDB, and Microsoft SQL Server. Scans over 8,000 hosts, tracks Java distributions and system owners, and sends notifications. Identified Oracle Java on 370+ hosts among more than 4,400 systems running Java.
- Developed a system to identify potentially unused VMs using 10+ indicators. Identified 99 candidates; 27 VMs were decommissioned, freeing 113 vCPUs and 353 GB of RAM.
- Maintained and extended Puppet configurations for the Linux fleet: migrated the infrastructure to Puppet 8, developed manifests and facts, and integrated Puppet with AWX and Git. Automated configuration checks, audit evidence collection, and enforcement of PCI DSS requirements for access controls and security settings.
- Collaborated with information security teams on vulnerability remediation and OS updates for systems within PCI DSS scope and other sensitive environments. Performed updates using preconfigured automation as well as other update methods.
- Automated server decommissioning and deployment of ClickHouse and Elasticsearch clusters, reducing manual work and the risk of human error.
- Operated AWS infrastructure and 50+ proxy servers across multiple regions; automated IP rotation and optimized EC2 instance selection. Configured access controls, network restrictions, and log forwarding to the SOC for AWS and AI infrastructure.