hirly

Apply with hirly

Sr. Staff Software Engineer, AI Agent Platform

Geico · New York City, NY · Palo Alto, CA +2

Upload your resume to see how well you match this job — free, in seconds, no account needed.

Your resume is used only to score it against this job. If you don't create an account, it is deleted within 24 hours.

Why Join GEICO? At GEICO, we offer a rewarding career where your ambitions are met with endless possibilities. Every day we honor our iconic brand by offering quality coverage to millions of customers and being there when they need us most. We thrive on relentless innovation to exceed our customers' expectations while making a real impact on local communities nationwide. Founded in 1936, GEICO is a member of the Berkshire Hathaway family of companies and one of the largest auto insurers in the United States. When you join our company, we want you to feel valued, supported, and proud to work here. That's why we offer the GEICO Pledge: Great Company, Great Culture, Great Rewards, and Great Careers. Sr. Staff Software Engineer, AI Agent Platform Why Join GEICO? GEICO is transforming how AI is built and deployed across the enterprise. As one of the largest insurers in the United States, we are investing heavily in next-generation AI platforms that empower more than 30,000 associates and enhance experiences for millions of customers. We are looking for a Sr. Staff Software Engineer with deep backend engineering and tech-lead experience to help build GEICO's enterprise AI Agent Platform. The team is building several core systems: an agent harness that runs agentic workflows durably in production, a retrieval-augmented generation (RAG) platform that grounds agents in enterprise knowledge, an insight engine that turns agent execution data and signals from related systems into live domain knowledge and context for both people and agents, an eval harness that measures agent quality and catches regressions, and a skills marketplace where teams publish and reuse agent capabilities. Each of them is a distributed systems problem at its core. We care about your track record designing and building durable, scalable systems as well as your experience in AI/ML. The Opportunity As a tech lead on the AI Agent Platform team, you will own the design and delivery of one or several of the core systems that power use-case-agnostic agentic workflows across the enterprise, from claims and underwriting to internal associate tools. This is platform work: generalized services and capabilities that many teams build on, not a single agent application. We welcome candidates from application, platform, or infrastructure backgrounds. What matters is that you have led the design of systems that stay correct, available, and operable under real production load, that you've been accountable for running them (on call, incidents, and all), and that you're comfortable working across the full backend stack. Hands-on experience building GenAI applications or agentic harnesses is a strong plus, and the foundation is systems judgment. What You Will Do Lead

  • Lead the technical design and delivery of major platform components end to end, from architecture and API design through rollout, operation, and iteration.
  • Drive design reviews, set engineering standards, and make sound tradeoff and build-vs-buy decisions for the platform.
  • Mentor engineers and raise the engineering bar across the team and the teams building on the platform.
  • Partner with product managers, data scientists, architects, and consuming teams to turn workflow needs into generalized platform capabilities. Build Depending on your strengths, you'll lead work in areas such as:
  • Agent harness: the durable execution engine for agent workflows, with state management, checkpointing, retries, idempotency, and recovery for long-running, multi-step, tool-calling workflows.
  • Agent harness integrations: services, APIs, and integration contracts (including protocols like MCP) that let many teams connect agents to internal services, data sources, and external systems consistently.
  • RAG platform: shared ingestion, chunking, embedding, indexing, and retrieval services that ground agents in enterprise documents and data, with access controls, freshness guarantees, and low-latency retrieval at scale.
  • Insight engine: pipelines and services that collect data from agent executions and related enterprise systems and turn it into continuously updated domain knowledge and agent context, served to people through dashboards and reports and to agents through retrieval and tool interfaces. Where the RAG platform grounds agents in relatively static documents, the insight engine captures what is changing: emerging patterns, outcomes, and operational signals.
  • Eval harness: high-throughput pipelines for running offline and online evaluations, simulating scenarios, capturing execution traces, detecting regressions, and storing large volumes of execution data, as a shared framework every team can use.
  • Skills marketplace: a multi-tenant service where teams publish, version, discover, and reuse agent skills and tools, backed by permissions, security review, and governance.
  • Reliability, isolation, and cost controls: sandboxing, rate limiting, quotas, SLOs, and capacity planning at enterprise scale.
  • Work across the stack as needed, from storage and data modeling to services, deployment infrastructure, and the developer-facing surfaces (SDKs, CLIs, portals) teams use to build on the platform. Operate
  • Own the production health of the systems you lead: define SLIs and SLOs, error budgets, and the dashboards and alerts that tell the team when something is wrong before users do.
  • Build observability in from the start, with metrics, structured logging, and distributed tracing that make failures in the agent harness and in complex, multi-step agent workflows quick to diagnose.
  • Design for scale and resilience through load testing, capacity planning, graceful degradation, and failure-mode analysis, so the platform holds up as adoption grows across the enterprise.
  • Participate in and help lead the on-call rotation, drive incident response, and run blameless postmortems that result in durable fixes rather than repeat incidents.
  • Champion operational excellence across the team through runbooks, alert hygiene, safe deployment practices (canaries, feature flags, rollbacks), and regular operational reviews. Minimum Qualifications
  • 10+ years of professional software engineering experience building and operating large-scale production backend systems.
  • Tech-lead experience: owned the design and delivery of significant systems involving multiple engineers or teams.
  • Strong proficiency in one or more of Python, Java, Go, or comparable languages.
  • Deep distributed systems fundamentals, including concurrency, consistency, fault tolerance, data modeling, API design, and scalability.
  • Experience operating production services with CI/CD, automated testing, and Kubernetes or other container platforms.
  • Hands-on operational experience owning the reliability of production systems, including observability, monitoring and alerting, on-call participation, incident response, and postmortems.
  • Proven experience scaling systems and improving reliability, such as defining and meeting SLOs, capacity planning, performance tuning, and eliminating recurring failure modes.
  • Background in application, platform, or infrastructure engineering; all are relevant. Preferred Qualifications
  • Experience with durable execution or workflow orchestration systems (e.g., Temporal, Restate, DBOS, Azure Durable Task, AWS Step Functions) or event-driven architectures (e.g., Kafka).
  • Experience building internal developer platforms, multi-tenant services, or SDKs used by many teams.
  • Experience building GenAI applications and agentic harnesses, using frontier and open-weight models (e.g., GPT, Claude, Gemini, Llama, Qwen) and frameworks such as LangGraph, Microsoft Agent Framework, OpenAI Agents SDK, or Claude Agent SDK.
  • Experience building RAG systems at scale, including ingestion pipelines, hybrid retrieval (dense plus keyword search), reranking, and permission-aware retrieval, using engines such as Azure AI Search, OpenSearch, or Ves
Apply: Sr. Staff Software Engineer, AI Agent Platform at Geico