Role: Principal NVIDIA AI Factory Solutions Architect
Location: Bay Area, CA
FTE Only
Mandatory - Nvidia tech stack - AI factory SME
ROLE OVERVIEW:
Our AI Services and Transformation Unit is seeking a hands-on Technology & Solution Architect who lives at the intersection of NVIDIA's AI Factory stack and enterprise client delivery. You are not a generalist - you have personally co- designed, configured, and validated AI Factory infrastructure using NVIDIA hardware and software. As GSI partner practitioner, you will work directly with clients and the NVIDIA partner engineering team to architect, co-design, and bring to production end-to-end AI Factory solutions - from GPU cluster topology and NVIDIA DSX validated designs to Agentic AI workload onboarding and token economics benchmarking.
WHAT YOU WILL DO:
AI Factory Co-Design & Client Advisory
Lead client-facing AI Factory co-design engagements: assess existing infrastructure, define target AI Factory architecture using NVIDIA Enterprise AI Factory Validated Design and Enterprise Reference Architectures, and
produce deployment-ready technical blueprints.
Advise clients on full NVIDIA AI Factory stack selection - compute (HGX B200/GB200/GB300, MGX, NVL72), networking (InfiniBand NDR/XDR, Spectrum-X), storage, and the NVIDIA AI software layer - matched to their Agentic AI, Physical AI, or HPC workload profile.
Apply NVIDIA DSX co-designed reference framework to architect modular, gigawatt-scalable AI Factories; use Omniverse DSX digital twin blueprints to simulate and validate designs pre-deployment.
Guide clients on cooling strategy for high-density GPU deployments (40 60+ kW/rack): Direct Liquid Cooling (DLC), immersion cooling, and RDHx - translating physics into procurement specifications and facility requirements.
NVIDIA GSI Partner Delivery
Function as unit practitioner-level interface with NVIDIA partner engineering - co-developing client solutions, navigating the GSI validated design process, and maintaining NVIDIA-Certified System configurations across client deployments.
Deploy and configure the NVIDIA AI Enterprise software suite (NIM microservices, NeMo, Nemotron, Dynamo, RAPIDS, Triton, Run:ai, Mission Control) on client AI Factory infrastructure.
Execute Agentic AI workload onboarding onto unified AI Factory platforms: implement NVIDIA AI Blueprint for RAG, configure cuOpt for operational AI, and deploy Kubernetes-native GPU orchestration using Run:ai and NVIDIA GPU/Network Operators.
Run GenAI-Perf and MLPerf benchmarks to validate AI Factory delivery quality; present token throughput (tokens/sec, tokens/watt, tokens/dollar) performance against client SLAs and competitive benchmarks. Technical Solutioning & Pre-Sales Support
Support pre-sales on AI Factory pursuits: build detailed BoMs, architecture diagrams, and NVIDIA stack solution documents for client proposals and RFP responses.
Quantify business value of AI Factory deployments - translate token economics, GPU utilization rates, and inference latency improvements into client ROI models.
Contribute to unit AI Factory IP: delivery accelerators, validated configuration templates, benchmark frameworks, and reusable reference architecture assets.
QUALIFICATIONS & EXPERIENCE
Experience
15+ years in technology consulting or infrastructure engineering; 4+ years specifically in AI Data Center, HPC, or GPU infrastructure delivery - hands-on build, configuration, and operate/manage is required.
Direct, verifiable experience working within or alongside NVIDIA's GSI or partner engineering program - this is non-
negotiable.
Proven track record deploying NVIDIA AI Factory components in production: you have racked, cabled, and configured DGX/HGX/MGX systems and validated NVIDIA-Certified configurations - not supervised a team that did.
Experience in client-facing consulting or advisory roles with technical solutioning and proposal ownership.
Must-Have Depth
Compute
HGX B200/GB200, GB300 NVL72, MGX, DGX SuperPOD, RTX PRO Server
Networking
InfiniBand NDR/XDR, Spectrum-X, NVLink 4/5, ConnectX-8, BlueField-3
AI Software
NIM, NeMo, Nemotron, Dynamo, RAPIDS, Triton, NVIDIA AI Enterprise, CUDA-X
Ops & Orch.
Mission Control, Run:ai, GPU Operator, NGC, Kubernetes, GenAI- Perf, MLPerf
Build Frameworks
DSX Ref Design, Enterprise AI Factory Validated Design, Omniverse DSX Digital Twin
Data & Agents
AI Blueprint for RAG, cuOpt, MIG, Confidential Compute, NVMe-oF storage
AI Workloads: LLM training/fine-tuning, Agentic AI deployment, inference optimization (FP4/FP8, speculative decoding), token economics benchmarking.
Education & Certifications:
BS/MS in Computer Science, Computer Architecture, Electrical Engineering, or equivalent applied engineering background.
NVIDIA NCP-AI or NCP-DS certification strongly preferred. CDCE or equivalent data center credential is a plus.
...What You Will Do We are looking for a highly motivated Full Stack Software Engineer to join our engineering team and contribute to the development of next-generation web applications, UI frameworks, and developer platforms. The ideal candidate should have strong experience...
..., X and YouTube. Job Description This role is field-based and candidates should live within a reasonable distance... ...,CA San Francisco, CA San Jose, CA The Field Reimbursement Manager (FRM) shares expertise and educates on Medical and Pharmacy...
Requisition ID: 32032 ArcelorMittal Calvert uses the most innovative technology to create the steel that tomorrows world will be made of. As part of a global organization, every day over 190,000 of our talented people, located in over 60 countries, push the boundaries...
...Job Title: Graphic Design Intern In-House Creative Team Reports to : Art Director Compensation: Unpaid; academic credit supported... ...and point-of-sale materials Prepare files for print production Participate in creative reviews and feedback sessions...
...Curri is looking for cargo van and sprinter van owner/operators for dedicated route work in the Potsdam, NY area. Curri is a delivery platform for construction and industrial supplies plumbing, electrical, HVAC, and building materials moving from suppliers to job sites...