Overview
What you'll be doing:
Support and contribute to SRE initiatives that improve reliability, scalability, and developer efficiency across enterprise systems.
Assist in building and maintaining distributed systems that power NVIDIA’s AI-powered enterprise products and services, learning modern architectural patterns along the way.
Help automate database operations, including provisioning, scaling, backup, and failover, for relational and vector database services.
Contribute to observability and monitoring efforts by building dashboards, alerts, and automation scripts to improve system performance and reliability.
Participate in incident response processes, learning to triage issues, reduce mean time to resolution (MTTR), and contribute to post-incident reviews.
Collaborate with Cloud, Platform, Security, and AI/ML teams to support platform reliability and help implement SRE best practices.
Learn to operate and troubleshoot complex systems, including Kubernetes-based and cloud-native infrastructure, following established standards in system design and incident management.
Explore and adopt AI-assisted engineering practices, including coding agents and LLM-powered tooling, to accelerate day-to-day development workflows.
What we need to see:
BS degree in Computer Science or a related technical field (e.g., physics, mathematics), or equivalent practical experience.
Foundational proficiency in at least one programming language such as Python, TypeScript, JavaScript, or Go.
Basic understanding of cloud platforms (AWS, Azure, or GCP) and containerization technologies like Docker and Kubernetes.
Exposure to or coursework in infrastructure-as-code tools (e.g., Terraform, AWS CDK, CloudFormation) or willingness to learn.
Familiarity with Linux/Unix systems, networking fundamentals, and version control (Git).
Interest in observability concepts (logging, metrics, tracing) and tools such as OpenTelemetry, Prometheus, or Grafana.
Basic knowledge of relational databases (e.g., PostgreSQL, MySQL), understanding of SQL, indexing, and simple query optimization.
Strong problem-solving skills, curiosity, and a willingness to learn in a fast-paced, collaborative environment.
Good communication and teamwork skills, with the ability to ask the right questions and learn from senior engineers.
Ways to stand out from the crowd:
Personal projects, internships, or coursework involving cloud infrastructure, automation, or DevOps/SRE practices.
Contributions to open-source projects or active participation in hackathons, coding competitions, or technical communities.
Exposure to AI/ML concepts, e.g., building or deploying a simple ML model, experimenting with LLM APIs, or using AI-powered developer tools (Copilot, Cursor, etc.).
Hands-on experience with CI/CD pipelines, scripting for automation, or container orchestration (even in personal or academic projects).
A strong sense of ownership, curiosity, and initiative, you turn challenges into learning opportunities and aren’t afraid to dive into unfamiliar systems.
If an employer asks you to pay any kind of fee, please notify us immediately. Talentd does not charge any fee from applicants and we do not allow other companies to do so.
Key Skills
Related Tags
Education Requirements
- BS degree in Computer Science or a related technical field (e.g., physics, mathematics)
- or equivalent practical experience
Eligible Batch Years

NVIDIA
NVIDIA Corporation is a global leader in accelerated computing, renowned for its pioneering work in graphics processing units (GPUs) and artificial intelligence (AI). Founded in 1993, the company has grown into one of the most influential technology firms in the world, with over 26,000 employees as of 2024. NVIDIA's mission is to advance computing technology to solve the world's most challenging problems, enabling breakthroughs in fields such as gaming, professional visualization, data science, autonomous vehicles, and high-performance computing.
The company is best known for its GeForce GPU product line, which has revolutionized the gaming industry, and its CUDA platform, which has empowered developers to harness GPU acceleration for AI and scientific research. In recent years, NVIDIA has expanded its portfolio through strategic acquisitions, such as Mellanox Technologies and Arm (pending regulatory approval), and has made significant strides in AI infrastructure, cloud services, and semiconductor innovation. Its market position is strong, with a reputation for cutting-edge technology and leadership in AI-driven computing, underscored by record revenues and a dominant role in powering generative AI models and large-scale data centers.
Company Details
Connect With Us
Websitenvidia.com