hirly

Apply with hirly

Director of Product Management, Agentforce Voice Models

Salesforce · California - San Francisco · Washington - Bellevue

Upload your resume to see how well you match this job — free, in seconds, no account needed.

Your resume is used only to score it against this job. If you don't create an account, it is deleted within 24 hours.

Already have an account? Sign in to see your saved application

To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts. Job Category Product Job Details About Salesforce Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn’t a buzzword — it’s a way of life. The world of work as we know it is changing and we're looking for Trailblazers who are passionate about bettering business and the world through AI, driving innovation, and keeping Salesforce's core values at the heart of it all. Ready to level-up your career at the company leading workforce transformation in the agentic era? You’re in the right place! Agentforce is the future of AI, and you are the future of Salesforce. About the Role Agentforce is Salesforce's next-generation AI platform, delivering autonomous agents that reason, take action, and communicate naturally across every customer touchpoint. Voice is the fastest-growing interaction surface—from contact-center automation and field-service assistants to real-time sales coaching and multilingual global deployments. As Director of Agentforce Voice Models, you will own the full voice intelligence stack: Automatic Speech Recognition (ASR / STT), Text-to-Speech (TTS), Speech-to-Speech (S2S) end-to-end pipelines, speaker diarization, prosody modeling, and the language-coverage roadmap that makes Agentforce sound natural in every market we serve. You will partner with product, infrastructure, and go-to-market teams to set the bar for transcription accuracy, latency, and voice expressiveness at enterprise scale. What You'll Do Voice Model Strategy & Roadmap

  • Define and own the multi-year roadmap for Agentforce voice capabilities, spanning ASR/STT, TTS, S2S, and real-time voice agents.
  • Set accuracy, latency, and quality benchmarks (WER, MOS, RTF, DMOS) and drive the organization to meet them.
  • Evaluate build vs. buy vs. partner decisions for new voice model capabilities and maintain relationships with key academic and industry partners. ASR / STT (Automatic Speech Recognition)
  • Lead the development of production-grade ASR systems optimized for telephony, WebRTC, and device-side deployment.Word Error Rate (WER) improvement across noise conditions, accents, and domain-specific vocabularyStreaming and batch recognition pipelines with sub-200 ms first-token latency targetsCustom vocabulary and language model adaptation (hot-word boosting, domain LM interpolation)Punctuation restoration and inverse text normalization (ITN) for downstream NLU
  • Drive multilingual and code-switching ASR coverage across priority languages; govern the language onboarding process including data acquisition, model training, and acceptance testing. TTS (Text-to-Speech)
  • Own the neural TTS pipeline—voice cloning, persona design, SSML compliance, and real-time synthesis—for Agentforce agent personas.
  • Lead prosody research: intonation, rhythm, stress, and pause modeling that produces natural-sounding enterprise voices across conversational contexts.
  • Manage voice talent agreements, ethical AI review, and consent frameworks for synthetic voice creation.
  • Drive naturalness, expressiveness, and brand-consistency quality bars using subjective (MOS, CMOS) and objective (mel-cepstral distortion) evaluation frameworks. Speech-to-Speech (S2S) & Real-Time Voice Agents
  • Architect low-latency S2S pipelines that enable full-duplex conversational AI without the ASR→NLU→TTS handoff penalty.
  • Partner with the Agentforce Reasoning team to integrate voice understanding with agent action loops (tool calls, CRM lookups, escalation routing).
  • Establish interruption, barge-in, and turn-taking models appropriate for enterprise voice agents. Speaker Diarization & Voice Analytics
  • Deliver production speaker diarization ("who spoke when") for multi-party calls, enabling accurate per-speaker transcripts used in call coaching, compliance, and analytics.
  • Develop speaker verification and voice-biometric capabilities for secure agent authentication use cases.
  • Partner with the Einstein Analytics team to surface voice-derived signals (sentiment, engagement, talk-time ratios) in Salesforce dashboards. Language Coverage & Localization
  • Own the global language support roadmap; prioritize languages by customer demand, addressable market, and data availability.
  • Establish data governance, annotation, and quality-control pipelines for low-resource languages.
  • Work with regional Salesforce teams to validate dialect, accent, and cultural appropriateness of voice personas. Engineering Leadership & Team Building
  • Recruit, develop, and retain a world-class team of research engineers, applied scientists, and ML engineers (target team size: 20–30).
  • Set technical direction, drive architectural decisions, and maintain engineering excellence through code review culture, rigorous evaluation frameworks, and production incident reviews.
  • Build a culture of experimentation: rapid A/B testing, red-teaming for voice safety, and continuous model refreshes. Cross-Functional Partnership
  • Partner with Product Management to translate customer and regulatory requirements into model specifications and acceptance criteria.
  • Collaborate with Legal, Privacy, and Trust & Safety on voice AI ethics, speaker consent, deepfake-detection, and compliance (HIPAA, PCI, GDPR).
  • Represent Agentforce Voice at external conferences, standards bodies (W3C, IETF), and customer briefings. Who You Are Required Experience & Skills
  • 10+ years in speech/audio machine learning, with 4+ years in a senior leadership role managing teams of 10 or more engineers or scientists.
  • Deep hands-on expertise in at least two of: ASR (end-to-end or hybrid CTC/attention architectures), neural TTS (VITS, Voicebox, Matcha-TTS, or equivalent), or real-time speech processing pipelines.
  • Proven track record of shipping production voice models at scale (millions of minutes per day) with measurable accuracy and latency improvements.
  • Fluency in prosody modeling: pitch contour modeling, duration prediction, and expressiveness fine-tuning.
  • Experience with speaker diarization systems (clustering-based, end-to-end, or hybrid).
  • Strong understanding of multilingual and multi-dialect NLP/ASR challenges; experience with at least 5 languages in production.
  • Proficiency in Python; experience with PyTorch or JAX model training at scale; familiarity with cloud-based training infrastructure (AWS, GCP, or Azure).
  • Excellent written and verbal communication—able to translate complex model behavior into product and business narratives for executives. Preferred Qualifications
  • PhD in Computer Science, Electrical Engineering, Linguistics, or related field with a speech/audio focus, or equivalent industry experience.
  • Familiarity with streaming inference, model quantization (INT8/INT4), and on-device deployment for low-latency voice.
  • Experience with voice safety, watermarking, and deepfake/spoofing detection.
  • Background in telephony protocols (SIP, RTP, WebRTC) and enterprise contact-center platforms (Genesys, NICE, Avaya, Amazon Connect).
  • Published research or patents in speech processing, audio ML, or a closely related field.
  • Experience operating within regulated industries (financial services, healthcare) with voice data compliance requirements. Why Salesforce & Agentforce At Salesforce, we believe AI agents will redefine how businesses and people work. Agentforce is at the center of that shift—and Voice is its most human interface. You will:
  • Work at the intersection of frontier research and enterprise-grade reliability, with direct customer impact across 150,000+ organizations worldwide.
  • Have access to proprietary, enterprise-grade voice data, world-class compute, and a platform with distribution that no startup can match.
  • Operate with the autonomy of a startup wit
Apply: Director of Product Management, Agentforce Voice Models at Salesforce