AI Gateway & Cost Governance: Own the vision for a governed, model-agnostic gateway that all model and agent traffic routes through, with per-team and per-use-case cost attribution, flexible model routing, rate limiting, provider failover, and automatic spending caps and stop switches. Make model-swap and cost decisions changeable once at the gateway, not per service.
Evaluation & Quality Enforcement: Define the roadmap for the eval engine and the pass/fail gate that runs on it, eval-gated prompt management with versioning and rollback, and per-agent accuracy scoring — so quality regressions are caught before they reach users rather than surfacing downstream in business metrics.
Trust & Autonomy Ladder: Define and champion the trust-and-autonomy ladder — the thresholds, progression criteria, evidence, approvals, and rollback logic that govern when an agent earns more independence — so agentic adoption scales with accountability instead of governance gaps.
LLM & ML Serving Infrastructure: Own the production path self-hosted LLM /open-weight model serving and traditional ML with the goal to consolidate both ML and LLM workloads under one standardized observable platform.
Cost & ROI Telemetry: Define and instrument the metrics that track the platform health as well as business impact to the platform.
Execution And Leadership
Cross-Functional Partnership: Establish deep partnerships with product engineering, architecture and data science to drive adoption and make sure the platform reflects how teams actually build and serve models.
Technical Roadmap Management: Manage a complex platform backlog spanning parallel tracks and under real capacity constraints. Make and communicate authoritative sequencing and trade-off decisions with concise and clear communications.
Stakeholder Communication: Act as the primary interface between technical platform teams and business stakeholders, translating infrastructure investment into clear business outcomes
Qualification
Required
8+ years of experience in Product Management, with at least 4 years focused on infrastructure, platforms, ML/AI systems, or other technical products serving internal engineering or data-science customers.
Proven track record scaling technical platforms from inception through maturity for demanding internal customers.
Technical fluency across the modern AI/ML stack — model serving and inference, API gateways/proxies, evaluation and testing frameworks, and cloud/Kubernetes infrastructure — sufficient to “swim with the fishes” with engineers and data scientists on architecture trade-offs.
Demonstrated ability to use telemetry and cost data (model performance, eval results, adoption analytics, spend data) to diagnose bottlenecks and drive product strategy.
Preferred
Hands-on familiarity with LLM gateways/proxies (e.g., LiteLLM), evaluation and observability tooling (e.g., Langfuse), and model serving frameworks (e.g., KServe, SageMaker, Bedrock, Ray).
Experience with prompt management and versioning, LLM fine-tuning, or self-hosted/open-weight model serving in production.
Familiarity with agent governance concepts — guardrails, PII controls, prompt-injection protection, and autonomy/trust frameworks — and with the tension between agent velocity and cost control.
Experience operating a platform across competing tracks and stakeholders, sequencing revenue-critical stabilization work alongside a new strategic build-out.
Exceptional leadership skills with a history of influencing cross-functional teams without direct authority.