Principal Supercomputing Operations Software Engineer
Active / Open Until Filled
Apply NowPrincipal Supercomputing Operations Software Engineer
Active / Open Until Filled
Apply NowJob role insights
-
Date posted
September 1, 2026
-
Closing date
Not Disclosed (Open Until Filled)
-
Location
Redmond
-
Salary
USD142,800 - USD274,800 /year
-
Career level
Principal / Lead
-
Experience
6+ Years
Description
Microsoft Corporation is actively seeking an exceptional professional to join our organization as a Principal Supercomputing Operations Software Engineer within the Software Engineering division. Positioned at the core of Microsoft’s technological innovation, this role offers direct involvement in mission-critical engineering initiatives, planetary-scale cloud services, and next-generation product architectures.
Role Overview & Strategic Impact
As a Principal Supercomputing Operations Software Engineer based in Redmond, Washington, United States (Multiple US Locations Available), you will play an integral role in architecting, engineering, and delivering transformative technology solutions. Operating within Microsoft’s Software Engineering practice, your efforts will directly shape the scalability, performance, and security of software systems and platforms trusted by billions of users worldwide.
Microsoft’s enduring mission is to empower every person and every organization on the planet to achieve more. Team members thrive within an inclusive culture grounded in a growth mindset, cross-team empathy, and uncompromising customer focus. In this position, you will collaborate with industry-leading engineers, researchers, and product leaders to solve complex technical problems at unprecedented scale.
Technical Domain: Hyperscale Cloud Architecture & Frontier AI Integration
Operating at the nexus of planetary-scale cloud infrastructure, Microsoft Azure powers the foundation for enterprise digital transformation and generative AI breakthroughs. Engineers and architects in this domain design, deploy, and optimize highly resilient, multi-region cloud services, intelligent API pipelines, and secure data topologies. Leveraging deep integrations with Azure OpenAI, foundational LLMs, and distributed microservices, team members accelerate production-grade intelligence while upholding strict standards for reliability, zero-trust security, and operational latency.
Professionals in this specialization combine strong computer science fundamentals with practical, real-world systems engineering. Whether delivering distributed microservices, designing custom silicon interconnects, managing physical datacenter infrastructure, or scaling AI-powered customer workflows, you will operate within a mature engineering ecosystem supported by Microsoft’s global technical community.
Key Responsibilities & Scope of Work
In this capacity, your daily priorities and core technical contributions will center on the following key operational deliverables:
- Serve as the technical authority and DRI for InfiniBand and GPU interconnect fabric operations across large scale AI supercomputing environments, ensuring sustained GPU availability, training stability, and SLA compliance
- Lead and orchestrate complex, high severity fabric incidents end to end, including detection, triage, mitigation, recovery, and root cause analysis, making high impact decisions under ambiguity
- Perform deep, multi layer systems debugging across InfiniBand, Subnet Manager, GPU interconnect, PCIe, GPUs, firmware, drivers, and OS layers to identify true root causes at fleet scale
- Drive operational excellence and systemic prevention by identifying recurring failure patterns, defining reliability models and failure domains, and authoring authoritative TSGs, playbooks, and escalation frameworks adopted across teams
- Architect and drive automation, telemetry, diagnostics, and tooling that materially improve detection, observability, debuggability, and mean time to mitigation, raising the operational bar for interconnect fabrics across the platform
- OR equivalent experience.
- Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings: Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter.
- Maintain rigorous code quality, comprehensive automated test coverage, and security compliance throughout all phases of software delivery.
- Participate actively in technical design reviews, operational post-mortems, and architectural discussions across global engineering cohorts.
Candidate Profile & Experience Requirements
Candidates for the Principal Supercomputing Operations Software Engineer position are evaluated based on their practical experience, technical depth, and demonstrated track record of problem-solving:
- Professional Background: Minimum of 6 years (72 months) of verified experience directly relevant to software engineering or related technical specializations.
- Bachelor’s Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
- Bachelor’s Degree in Computer Science OR related technical field AND 10+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, OR Python OR Master’s Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
- OR Master’s Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
- 6+ years of experience operating large‑scale distributed systems, high‑performance computing (HPC), or artificial intelligence (AI) infrastructure in production environments Demonstrated ownership of mission‑critical production infrastructure with direct impact on service availability, GPU workloads, and customer SLAs Hands‑on experience operating and debugging interconnect fabrics supporting large‑scale compute workloads Strong Linux systems knowledge with experience debugging low‑level infrastructure issues across operating systems, drivers, and services Proven ability to reason across hardware, firmware, drivers, and software stacks to diagnose and resolve complex production issues
- Technical Competence: Demonstrated mastery of programming environments, distributed systems, modern cloud infrastructure, or specialized analytical methodologies relevant to the domain.
- Cross-Functional Leadership: Strong written and verbal communication skills, with an established ability to articulate complex technical architectures clearly to engineering peers and executive stakeholders.
- Culture & Values: Demonstrated alignment with Microsoft’s core values of respect, integrity, and accountability, alongside a proactive dedication to diversity and inclusion.
Work Environment, Workplace Schedule & Location Structure
This position is designated with an official location in Redmond, Washington, United States (Multiple US Locations Available). Work arrangements for this role are structured as Up to 100% Work from Home / Hybrid Flexibility (Across US Locations). Microsoft provides modern, accessible work environments equipped with state-of-the-art developer tooling, ergonomic facilities, and collaborative hybrid infrastructure.
Professional Growth, Technical Community & Continuous Learning
At Microsoft, career progression is treated as an ongoing, deliberate journey supported by structured enterprise resources. Team members within the Software Engineering organization are empowered to continuously expand their technical capabilities, explore emerging engineering paradigms, and pursue recognized professional credentials. Global technical communities and internal symposiums allow engineers to exchange architectural insights and collaborate across diverse product portfolios.
Whether advancing along specialized individual contributor tracks toward principal architecture or transitioning into technical leadership and people management, employees benefit from transparent career milestones, executive sponsorship, and comprehensive mentoring networks designed to nurture sustainable professional excellence.
Compensation, Total Rewards & Comprehensive Benefits
Microsoft offers a comprehensive, highly competitive total rewards framework designed to support employees’ personal, professional, and financial well-being:
- Legally Disclosed Base Salary Range: USD $142,800 – USD $274,800 per annum across typical US geographic zones (regional geographic differentials apply for high-cost zones such as the San Francisco Bay Area and New York City metropolitan areas).
- Incentive & Equity Rewards: Eligibility for annual performance-based cash bonuses and Microsoft restricted stock units (RSUs).
- Comprehensive Healthcare: Premium medical, dental, and vision coverage, including mental health support programs and wellness reimbursements.
- Retirement & Financial Security: Robust 401(k) / national retirement savings plans with generous company matching contributions, alongside financial planning resources.
- Time Off & Work-Life Harmony: Generous paid time off, comprehensive parental leave programs, public holiday allowances, and caregiver leave support.
- Growth & Learning: Unlimited access to Microsoft Learn, internal technical conferences, tuition assistance, and structured career mobility pathways across worldwide business units.
Application Process & Requisition Reference
This position is actively open and monitored by Microsoft Talent Acquisition. In alignment with corporate hiring policies, applications are accepted on an ongoing basis with no rigid pre-determined closing deadline until a qualified candidate is formally selected (Open Until Filled).
To submit your application, select the Apply Now button on this listing to navigate directly to the official Microsoft Careers requisition portal. When submitting your candidacy, reference Requisition ID 200028314. Ensure your resume and supporting documentation highlight your relevant technical deliverables, system achievements, and professional qualifications.
