Cisco
US - North Carolina - Triangle
View Company Profile /
<< Go Back
Own application-level reliability for AI-powered features by defining SLIs, SLOs, and error budgets, and building observability dashboards and automated diagnostic and remediation agents. Lead application incident response and reliability improvements, using operational analytics, Python tooling, and collaboration with application, data, and infrastructure teams.
Requirements: Requires at least 7 years of software engineering experience focused on reliability, observability, or production operations, plus a bachelor's or master's degree in a relevant technical discipline. Candidates need strong Python and SQL skills, production GCP experience with GKE, BigQuery, and Bigtable or an equivalent, and experience operating application-level SLI/SLO frameworks and debugging production applications.
Key Skills: Python, Kubernetes, Google Cloud Platform, BigQuery, Bigtable, SQL, Application Reliability, Observability, SLI/SLO Frameworks, LangGraph, AI Agent Engineering, Looker, Distributed Tracing, Incident Response, Automated Remediation, Generative AI
Benefits: Medical Insurance, Dental Insurance, Vision Insurance, 401(k) Matching, Paid Parental Leave, Short-Term Disability Coverage, Long-Term Disability Coverage, Life Insurance, Paid Holidays, Floating Holiday, Birthday Leave, Year-End Holiday Shutdown, Personal Wellness Days, Paid Vacation, Sick Leave, Family Emergency Leave, Volunteer Leave, Annual Bonuses, Restricted Stock Units
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX