AI Site Reliability Engineer

Lockheed Martin Corporation United States

Apply now Save · sign in Alert me to jobs like this
Salary $73k - $188k/yr
Posted 07 Oct 2026
Last seen 07 Oct 2026
Location United States
Lifecycle fresh
Grade B

This listing scored 79/100, which is a B. It lost the most ground on remote clarity.

Every figure above is arithmetic over the posting itself — its salary field, its text, its age, its tags and how many sources carry it. How the grades work →

Developer Mid level Full Time

Standard Job Description

Your Mission:

The AI Factory team is seeking a Site Reliability Engineer to improve the reliability of the platforms and services that support AI development and deployment. In this role, you will use software development and automation to identify operational problems, reduce repetitive work, and help teams deliver dependable systems.

You will work with engineers and other partners across the AI Factory to improve observability, incident response, release validation, and platform health. The role is broader than any single tool or product: you may contribute to reliability platforms such as Cluster Concierge, but your focus will be on solving reliability problems across the environment.

Key Responsibilities:

· Improve the scalability, resilience, and reliability of existing platform services and middleware to ensure they remain dependable as usage and demand grow.

· Develop and maintain automation and operational tooling that improve platform reliability and reduce recurring manual work.

· Help teams investigate incidents, identify contributing factors, and implement fixes that prevent repeat issues.

· Build and improve dashboards, alerts, and other observability capabilities using metrics, logs, and traces.

· Create automated tests and validation workflows for upgrades, releases, and changes to platform services.

· Contribute to CI/CD and GitOps workflows that support consistent, reliable deployments.

· Assess platform health, document findings, and work with partner teams on practical reliability improvements.

· Participate in design reviews, code reviews, testing, and incident reviews.

· Contribute to reliability improvements for AI Factory services, including AIF Up and tools that support health validation and investigation.

Responsible for autonomy hardware and software system integration and testing including verification and validation.Translates customer requirements into product and systems specifications while addressing technical, schedule, and cost considerations; Establishes functional and technical specifications and standards for autonomous systems; Determines sensing hardware components; Recommends and selects appropriate control systems; Integrates and optimizes the autonomous compute and sensing hardware and software; Solves hardware/software interface problems; Develops plan(s) to integrate autonomous functionality into product(s) and platform(s) and other system(s); Develops test plans for validation and verification and procedures for test and evaluation requirements in collaboration with designers and developers; Performs integration testing and coordinates subsystem and/or system testing activities for autonomous programs; Documents and conducts analysis of test results and recommends fixes to software, hardware components, subsystems and systems; Interfaces with other teams involved the development lifecycle for perception

Basic Qualifications

· Experience designing, developing, and maintaining production software or automation using Go, Python, or a comparable language.

· Experience operating or engineering Kubernetes-based platforms, including troubleshooting complex service or infrastructure issues.

· Experience building or improving CI/CD, GitOps, infrastructure automation, or deployment workflows.

· Experience using observability data—including metrics, logs, or traces—to diagnose problems and improve system health.

Desired Skills

· Experience with OpenShift, GitLab CI/CD, Argo CD, Argo Rollouts, or similar GitOps and progressive-delivery tooling.

· Experience improving incident response, reducing operational toil, or defining actionable service-health measures.

· Familiarity with Prometheus, Grafana, OpenTelemetry, or comparable observability tools.

· Familiarity with designing automated reliability tests, upgrade validation, resilience tests, or failure-mode analyses.

· Familiarity with AI/ML platforms, GPU-based infrastructure, or deployments in disconnected environments.

· Strong oral and written communication skills, and ability to collaborate with cross-functional partners

· Creative and resourceful when it comes to problem-solving

· Ability to work with internal stakeholders to collect feedback, prioritize tasks, and manage the engineering backlog

· Self-motivated, self-directed, and the ability to thrive in a fast-paced environment in an industry that constantly changes

Pay Information

GeoZone Definition: GeoZones are geographic groupings created by Lockheed Martin to align compensation ranges with regional labor markets and cost-of-labor differences across the United States. Locations are assigned a Geo Zone based on the primary work location of the role.

At Lockheed Martin, we know mission success starts with taking care of our people. Our Total Rewards program is designed to attract top talent, support your well-being, and help you grow—both professionally and personally.

The salary range for this position is as listed on the requisition. Please note that the salary information listed is a general guideline only.
Lockheed Martin considers factors such as (but not limited to) scope and responsibilities of the position, candidate's work experience, education/ training, key skills as well as market(work location) and business considerations when extending an offer.

Benefits offered: Medical, Dental, Vision, Flexible work arrangements and schedules (e.g., 4x10), 401(k) match, Paid time off, Holidays, Parental Leave, EAP, Flexible Spending Accounts, Education Assistance, Life Insurance, Short-Term Disability, and Long-Term Disability.

Originally posted on Himalayas

Apply for this role Opens himalayas.app — the link as listed; we have not yet verified it is the employer's own page

Quick question · anonymous · one tap

Would you apply to this job?

Answer to see what other job seekers said.

Keep looking

Similar remote roles, still open

See every "AI Site Reliability" role →

Your turn · no account needed

Help the next applicant

You may know something about this listing that we cannot see from here. One tap. No account needed. Signed-in reports earn points once the evidence agrees with you.

I know what it pays

What you were offered, quoted in an interview, or paid in this role. A range is fine.

Sign in with Google to earn points for reports — 100 confirmed points buy a week of Early Access.

Where this listing came from

  1. 07 Oct 2026 Himalayas first sighting

Seen on 1 board over 0 days.