Intern - Data Scientist
Summary
Build ML models and data pipelines to analyze terabyte-to-petabyte semiconductor manufacturing datasets, improving yield and detecting anomalies using Python, Spark, and SQL.
Our vision is to transform how the world uses information to enrich life for all. Join an inclusive team passionate about one thing: using their expertise in the relentless pursuit of innovation for customers and partners. The solutions we build help make everything from virtual reality experiences to breakthroughs in neural networks possible. We do it all while committing to integrity, sustainability, and giving back to our communities. Because doing so can fuel the very innovation we are pursuing. Project Description This project focuses on applying data science, machine learning, and statistical modeling to large and diverse datasets generated from highly automated semiconductor manufacturing operations. The intern will work closely with data scientists and engineers to extract, process, and analyze terabytes to petabytes of structured and unstructured data. Objective Of The Project
- Develop analytical and machine learning models to improve manufacturing performance and operational efficiency
- Build data pipelines for extraction, cleansing, outlier detection, and anomaly identification
- Apply supervised, unsupervised, and semi-supervised learning techniques on real production datasets
- Support digital transformation and automation efforts in a manufacturing environment
- Data Scientist
- Machine Learning Engineer
- Data Engineer
- Smart Manufacturing and Analytics Engineer
- Extract and manipulate data using SQL and large-scale query systems
- Clean and prepare data using statistical and programming techniques
- Build machine learning models using tools such as Python or R
- Work with distributed computing systems including Spark and Hadoop-based tools
- Support automation and analysis workflows for manufacturing datasets such as time-series and images
- Collaborate with multi-functional teams to translate technical findings into actionable insights
- Exposure to highvolume semiconductor manufacturing data
- Hands-on experience with machine learning, deep learning, and data preparation workflows
- Familiarity with large-scale data platforms (Spark, Hadoop, Teradata)
- Experience with Manufacturing Execution Systems
- Opportunity to apply statistical modeling, feature extraction, and algorithm development
- Development of strong communication and problem-solving skills through cross-team collaboration
- Completed data pipelines or analytical scripts
- Machine learning models or prototypes with documented performance metrics
- Clear and concise technical reports summarizing findings
- Dashboards, visualizations, or automated analytics tools where applicable
- Supports data-driven decision-making in manufacturing
- Improves production efficiency, yield, and process stability
- Enhances predictive capabilities for production anomalies and variations
- Contributes to long-term digital transformation and automation initiatives
- Strong interest in building a career in data science
- Knowledge in statistical modeling and machine learning
- Experience with Python or R
- Ability to write efficient and maintainable code
- Familiarity with SQL
- Interest or experience in distributed computing systems (Spark, Hadoop)
- Strong communication and documentation skills
- Experience with time-series, image data, or semi-supervised learning
- Knowledge of JavaScript or data visualization (Tableau)
- Exposure to manufacturing systems or semiconductor processes
- Experience with TensorFlow or other deep learning frameworks
- Mathematics
- Data Science
- Computer Science
- Physics
- Statistics
- Engineering programs with analytics emphasis