
Leading supplier of end-to-end high speed Ethernet and InfiniBand intelligent interconnect solutions and services.

Leading supplier of end-to-end high speed Ethernet and InfiniBand intelligent interconnect solutions and services.
Company: NVIDIA
Core focus: Accelerated computing and AI infrastructure (GPUs, data center and AI systems, edge/robotics, automotive)
Founded: April 1993
Public company / Investor relations: Maintains an active investor relations portal with quarterly results and presentations
High-performance graphics, machine learning/AI acceleration, and full-stack AI infrastructure for enterprise, cloud and edge use cases.
1993
Semiconductors / AI infrastructure
Historical financing and IPO-related information are documented on financial profiles; NVIDIA also deploys strategic investments in startups (examples reported)
$2B
Reported $2B investment to support CoreWeave's AI compute expansion
NVIDIA’s AI Infrastructure organization is seeking a Senior AI Observability Engineer to help architect and implement distributed observability systems for AI and HPC clusters. We serve and collaborate directly with NVIDIA’s rapidly growing AI, HW, and SW engineering and research teams across the company. You will be working with a team of dedicated engineers on systems for data collection, aggregation, enrichment, storage, retrieval, and visualization to spectacularly improve efficiency, performance, and productivity of AI and HPC workloads. You will develop, deploy, and operate observability solutions for multiple compute clusters around the world. What You’ll Be Doing: * Collaborate with AI, HW, SW engineering and research teams to deliver observability solutions that meet their needs in AI/HPC clusters. * Develop, test, and deploy data collectors, pipelines, visualization and retrieval services. * Build a self-serve platform * Define data collection and retention policies to balance network bandwidth, system load, and storage capacity costs with data analysis requirements. * Work in a diverse team to provide operational and strategic data to empower our engineers and researchers to improve performance, productivity, and efficiency. * Continuously improve quality, workloads, and processes through better observability. What We Need to See: * Experience developing large scale, distributed observability systems. * Ability to collaborate with data scientists, researchers, and engineering teams to identify high value data for collection and analysis. * Experience with turning raw data into actionable reports * Experience with observability platforms such as Apache Spark, Elastic/Open Search, Grafana, Prometheus, and other similar open-source tools * Python programming experience and use of API calls * Passion for improving the productivity of others * Excellent planning and interpersonal skills * Flexibility/adaptability working in a dynamic environment with changing requirements * MS (preferred) or BS in Computer Science, Electrical Engineering, or related field (or equivalent experience) * 8+ yrs of proven experience. Ways To Stand Out from The Crowd: * Practical experience in machine learning, deep learning, open-source software, infrastructure technologies, and GPU technology. * Prior experience in infrastructure software, production application software development, software development, release and support methodology and DevOps * Experience in the management of datacenters and large-scale distributed computing * Experience working with AI researchers and/or EDA developers * Consistent track record of driving process improvements and measuring efficiency and a passion for sharing knowledge and experience driving complex projects end-to-end. The base salary range is 184,000 USD - 356,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis. NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Your next opportunity is in here somewhere. Sign up to explore 52,000+ startups and their open roles. No spam. No gamification. Just jobs.
52,000+
Startups
66,000+
Open Roles
1,400+
New This Week