Position Details: Sr. Infrastructure/Platform QA Engineer
Description:
Job Title: Senior Infrastructure/Platform QA Engineer
About the Role
We are looking for a Senior QA Engineer to own quality across our data protection and automated recovery platform. This is not a conventional QA role. You will be embedded in our engineering sprints, responsible not only for validating features but for deeply diagnosing failures across a complex, multi-layered stack spanning cloud-native microservices, enterprise backup products, hypervisor infrastructure, and Linux-based workloads.
You will own and maintain a live QA lab environment that mirrors real customer deployments, execute and automate test scenarios end-to-end, and be the person who can look at a failed test recovery and trace the root cause — whether it lives in the API payload, the backup catalog, the storage layer, or the network configuration.
What You'll Do
Quality Engineering
- Participate in sprint planning and design test strategies for new features before development begins
- Execute functional, integration, regression, and exploratory testing across the full product stack
- Go beyond surface-level failure reporting — diagnose why failures occur by examining logs, API payloads, infrastructure state, and component interactions
- Write and maintain automated test scripts (UI, API, and integration-level)
- File detailed, reproducible bug reports with root cause analysis, not just reproduction steps
Infrastructure & Lab Ownership
- Build, maintain, and evolve a QA lab environment that includes: VMware vSphere and Microsoft Hyper-V clusters, multiple backup product instances (IBM Storage Protect/TSM, Veeam, Cohesity, Rubrik, Zerto, NovaStor, Storage Protect for VE), a configured scanning VM with Ansible, AWX, Trend Micro, and ClamAV running as containers, and realistic backup schedules and datasets
- Simulate customer environment configurations, including edge cases and failure modes (storage offline, insufficient space, misconfigured DNS, expired credentials, etc.)
- Maintain environment stability so that sprint testing is never blocked by lab issues
Workflows Validation
- Validate test recovery workflows end-to-end for VMware and Hyper-V VMs across all supported backup products
- Diagnose recovery failures at the appropriate layer: missing/incorrect backup snapshot, wrong datetime passed in API call, target storage offline or undersized, permission issues, network routing failures, etc.
- Validate workflows: datastore mount/unmount operations, Ansible playbook execution, AWX job runs, scan results, and post-scan reporting
- Diagnose failures: DNS misconfiguration in scanning VM, failed datastore mount, stopped containers, misconfigured Trend Micro manager, Ansible connectivity issues, etc.
Collaboration
- Work closely with developers to reproduce issues, validate fixes, and prevent regressions
- Contribute to test documentation, runbooks, and QA environment setup guides
- Provide sprint-end quality summaries and risk assessments to the team
What You'll Need
Must-Have
- Solid understanding of VMware vSphere and/or Microsoft Hyper-V administration — you should be comfortable in vCenter and understand how snapshots, datastores, and VM recovery work at an operational level
- Hands-on experience with at least one enterprise backup product (IBM Storage Protect, Veeam, Cohesity, Rubrik, or similar)
- Knowledge of networking fundamentals: DNS, routing, firewall rules — enough to identify misconfiguration as a failure cause
- Experience with SSH-based remote command execution and Linux troubleshooting
- Proficiency in reading and interpreting logs, API responses (REST/JSON), and command-line output to diagnose failures
- Working knowledge of Windows and Linux systems administration
- Familiarity with PowerShell and/or Bash scripting
- Experience testing REST APIs (Postman, curl, or equivalent)
- Strong written communication — able to write clear, detailed bug reports with root cause analysis
Strong Advantage
- Experience with Ansible and AWX (or similar automation platforms)
- Familiarity with containerized workloads (Docker) and diagnosing container-level issues
- Experience with .NET application stacks or reading .NET/Windows service logs
- Familiarity with agile/scrum development workflows
Nice to Have
- Experience with any of: Zerto, NovaStor, Cohesity, Rubrik, IBM Storage Protect for Virtual Environments
- Any test automation framework experience (Selenium, Playwright, Cypress for UI; RestSharp, pytest for API)
Apply to Position