Provide senior SRE expertise to improve reliability, scalability, performance, and resilience of a cloud-hosted geospatial platform. Design monitoring/observability, automate deployments, support incident response, optimize capacity and performance, and collaborate across DevSecOps, Kubernetes, database, and support teams in a mission-focused DoD environment.
The Site Reliability Engineer (SRE) / Subject Matter Expert (SME) – Computer Systems Engineer/Architect will provide senior-level reach-back expertise to support the reliability, scalability, performance, and operational resilience of the GEOMAP platform in secure cloud environments. This role focuses on improving service availability, monitoring, incident response, automation, and production stability across cloud-hosted and containerized systems supporting mission-critical geospatial capabilities for the U.S. Air Force.
The Site Reliability Engineer will collaborate across development, DevSecOps, cloud, database, testing, and support teams to identify systemic issues, reduce operational risk, and implement engineering solutions that improve long-term platform reliability.
*This position is contingent upon contract award.*
Responsibilities
- Provide senior-level engineering support to improve reliability, availability, performance, and maintainability of GEOMAP cloud-hosted systems and services.
- Analyze production issues, recurring incidents, and operational trends to identify root causes and recommend durable corrective actions.
- Support the design and implementation of monitoring, alerting, logging, and observability solutions across applications, infrastructure, and containerized services.
- Develop and recommend automation approaches that reduce manual effort, improve deployment consistency, and increase system resilience.
- Partner with software engineers, DevSecOps engineers, Kubernetes engineers, database engineers, and production support personnel to improve service health and release readiness.
- Support incident response, problem management, service restoration, and post-incident reviews for high-priority operational issues.
- Evaluate system performance, capacity, and scalability needs and provide recommendations for optimization and operational risk reduction.
- Assist in defining service reliability objectives, operational metrics, and support models for sustained mission operations.
- Contribute to infrastructure and platform engineering efforts involving cloud environments, CI/CD pipelines, container orchestration, and secure deployment patterns.
- Support architecture reviews, technical assessments, and engineering analyses related to reliability, recoverability, and production operations.
- Develop or refine runbooks, standard operating procedures, reliability engineering practices, and technical documentation.
- Provide reach-back support for surge requirements, complex production investigations, and priority modernization or stabilization efforts as directed.
- Performs other related duties as assigned.
Qualifications
- Active Secret clearance required.
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field; Master’s degree preferred.
- Minimum of 8 years of experience supporting enterprise systems, cloud platforms, site reliability engineering, production engineering, systems engineering, or related technical roles.
- Experience supporting AWS environments, including monitoring, performance tuning, troubleshooting, incident response, and operational sustainment.
- Experience with Linux administration, scripting, and troubleshooting distributed applications in production environments.
- Experience with containerized systems and orchestration platforms such as Kubernetes.
- Experience supporting CI/CD pipelines, release automation, infrastructure-as-code, and operational reliability in Agile or DevSecOps environments.
- Experience with monitoring, logging, and alerting tools used to support enterprise application performance and infrastructure visibility.
- Strong analytical, troubleshooting, documentation, and communication skills, with the ability to translate operational issues into engineering improvements.
- Ability to work effectively across cross-functional teams in a mission-focused DoD environment.
Preferred
- Experience supporting AWS Cloud One or other secure federal cloud environments.
- Experience supporting geospatial or Esri-based platforms, including ArcGIS Enterprise or related technologies.
- Familiarity with service reliability practices such as SLIs, SLOs, error budgets, incident postmortems, and capacity planning.
- Experience with Risk Management Framework (RMF), STIG compliance, vulnerability remediation, and secure system hardening practices.
- AWS, Kubernetes, or other relevant cloud or reliability engineering certifications.
- Experience supporting technical refresh, platform modernization, or high-availability design initiatives in enterprise environments.
About Us
Diné Development Corporation (DDC) is a Navajo Nation owned family of companies that provides government agencies and commercial organizations with high-quality IT, professional, environmental, and research and development services. DDC is dedicated to empowering the Navajo Nation and communities we serve.
Diné Development Corporation (DDC) is a Navajo Nation owned family of companies that provides government agencies and commercial organizations with high-quality IT, professional, environmental, and research and development services. DDC is dedicated to empowering the Navajo Nation and communities we serve.
Benefits
Eligible full-time employees receive a comprehensive benefits package, including medical, dental, vision, life and disability coverage, retirement savings with company match, paid time off, voluntary supplemental benefits, and access to an employee assistance program. The package also includes educational assistance, with tuition reimbursement.
Eligible full-time employees receive a comprehensive benefits package, including medical, dental, vision, life and disability coverage, retirement savings with company match, paid time off, voluntary supplemental benefits, and access to an employee assistance program. The package also includes educational assistance, with tuition reimbursement.
EEO Statement
This contractor and subcontractor shall abide by the requirements of 41 CFR 60-1.4(a), 60-300.5(a), and 60-741.5(a). These regulations prohibit discrimination against qualified individuals based on their status as protected veterans or individuals with disabilities, and prohibit discrimination against all individuals based on their race, color, religion, sex, sexual orientation, gender identity, national origin, or for inquiring about, discussing, or disclosing information about compensation, or any other basis prohibited by law. We participate in E-Verify.
This contractor and subcontractor shall abide by the requirements of 41 CFR 60-1.4(a), 60-300.5(a), and 60-741.5(a). These regulations prohibit discrimination against qualified individuals based on their status as protected veterans or individuals with disabilities, and prohibit discrimination against all individuals based on their race, color, religion, sex, sexual orientation, gender identity, national origin, or for inquiring about, discussing, or disclosing information about compensation, or any other basis prohibited by law. We participate in E-Verify.
Similar Jobs
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads the architecture, modernization, optimization, reliability, and performance of enterprise IBM z/OS Db2 systems. Provides technical guidance across development and operations teams, automates DDL/DML processes, supports resilience and business continuity testing, resolves Db2 incidents, analyzes performance telemetry, improves SQL and database design, and partners with architects and stakeholders on enterprise technology strategy and hybrid-cloud modernization.
Top Skills:
AnsibleCloud IntegrationDdlDevOpsDmlIbm Db2Ibm Z/OsOpenshiftPythonRed Hat AnsibleRmfSmfSQLZlinux
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Designs and operates secure, reliable Azure cloud platforms using Terraform, GitHub Actions, containers, and automation. Responsibilities include CI/CD, observability, incident response, platform security, vulnerability remediation, disaster recovery, infrastructure troubleshooting, and SRE practices. The role supports production workloads, improves reliability and delivery processes, participates in on-call activities, and mentors engineers while partnering across development, security, architecture, and operations teams.
Top Skills:
BashCi/CdCloud SecurityDockerGitGithub ActionsGitopsInfrastructure As CodeKubernetesAzureObservabilityPowershellPythonTerraform
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads architecture, modernization, optimization, and reliability initiatives for mainframe CICS, MQ, and z/OS Connect environments. Provides technical direction across development and operations teams, establishes governance and change processes, tunes performance using telemetry, resolves incidents, and develops modernization roadmaps. Collaborates with stakeholders and enterprise architects to deliver secure, scalable, high-availability solutions while evaluating automation, cloud integration, and AI technologies.
Top Skills:
AnsibleCicsCobolDevOpsIbm MqIbm Z/OsOpenshiftPythonRed Hat Ansible Automation PlatformZ/Os ConnectZlinux
What you need to know about the Boston Tech Scene
Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.
Key Facts About Boston Tech
- Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
- Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
- Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
- Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories


