Job Description
We’re forming a team of innovators to roll out and enhance AI inference solutions at scale, demonstrating NVIDIA’s GPU technology and Kubernetes. As a Solutions Architect focused on inference, you’ll collaborate closely with our engineering, DevOps, and customers to develop enterprise AI solutions. Together, we'll deliver generative AI to production!
What you'll be doing:
+ Build inference pipelines with tools like NVIDIA Dynamo, distributing tasks among GPU workers to improve efficiency.
+ Collaborate with DevOps teams to orchestrate disaggregated inference using Kubernetes for complex workloads.
+ Accelerate inference pipelines using TensorRT-LLM, vLLM, SGLang, and other backends to ensure seamless integration with disaggregated inference.
+ Provide mentorship and technical leadership to customers and internal teams, guiding them through the deployment of disaggregated inference systems and resolving complex issues.
What we need to see:...
What you'll be doing:
+ Build inference pipelines with tools like NVIDIA Dynamo, distributing tasks among GPU workers to improve efficiency.
+ Collaborate with DevOps teams to orchestrate disaggregated inference using Kubernetes for complex workloads.
+ Accelerate inference pipelines using TensorRT-LLM, vLLM, SGLang, and other backends to ensure seamless integration with disaggregated inference.
+ Provide mentorship and technical leadership to customers and internal teams, guiding them through the deployment of disaggregated inference systems and resolving complex issues.
What we need to see:...
Ready to Apply?
Submit your application today and join our talented team at NVIDIA.
Submit ApplicationJob Details
- Location Santa Clara, CA
- Job Type Full-time
- Category other-general
- Posted Date June 06, 2026
- Application Deadline June 11, 2026