Job Description:
AI / Machine Learning Engineer
Position Overview
We are seeking a high-impact AI/ML Engineer to build intelligent data products that transform complex, high-volume engineering information into trusted, actionable insights. This individual will work across applied machine learning, generative AI, data platforms, and cloud engineering to deliver production-ready solutions supporting search, traceability, analytics, and decision-making.
This role is ideal for an engineer who can move seamlessly from architecture and design to implementation and operational ownership, while solving ambiguous problems where data quality, scalability, and reliability are critical.
Key Responsibilities
AI, Machine Learning & Data Products
- Architect, build, and operate reliable data products that ingest and transform structured and unstructured data at enterprise scale.
- Develop production-grade Retrieval-Augmented Generation (RAG) systems that combine semantic search, structured data, and grounded responses for engineering use cases.
- Design and implement agentic AI workflows capable of:
- Decomposing complex questions
- Selecting appropriate data sources and tools
- Validating results
- Returning explainable answers with citations
- Develop and evaluate embedding models, document understanding solutions, and multimodal inference workflows while balancing quality, scalability, latency, and cost.
- Deliver analytics and decision-support tools that enable engineers, program teams, and leadership to make informed decisions.
Data Engineering & Platform Development
- Build scalable data pipelines and processing systems for large, evolving datasets.
- Create resilient orchestration patterns for batch and near real-time workloads.
- Develop restartable and checkpoint-enabled processing architectures that maintain data integrity during long-running or partially failed workloads.
- Establish data quality, lineage, validation, and governance standards.
Cloud & Software Engineering
- Lead cloud architecture initiatives, including containerization, infrastructure-as-code, and CI/CD deployment practices.
- Design secure, repeatable deployment processes across multiple environments.
- Implement observability, alerting, monitoring, and operational runbooks.
- Ensure system reliability through incident investigation, performance tuning, root-cause analysis, and capacity planning.
Required Skills
- Python
- SQL
- Artificial Intelligence / Machine Learning
- Expert Systems
- Google Cloud Platform (GCP)
- API Development & Integration
- Software Testing
- Data Analysis
Preferred Skills
- Data & Analytics Dashboards
- Data Collection & Data Acquisition
- Data Integrity
- Data Conversion
- Java
Required Experience
- 5+ years of experience building and operating production software, data, or machine learning systems.
- Strong proficiency in Python and SQL.
- Experience with cloud platforms, managed data services, object storage, containers, and distributed computing workloads.
- Experience designing and operating scalable data pipelines and distributed processing systems.
- Hands-on experience applying Large Language Models (LLMs) in production environments, including:
- Prompt engineering
- Structured outputs
- Tool integration
- Model evaluation
- Production monitoring
- Deep understanding of:
- Embeddings
- Vector databases and retrieval
- RAG architectures
- LLM limitations and mitigation techniques
- Answer quality and faithfulness optimization
- Experience with software engineering best practices, including:
- Testing
- Code reviews
- Version control
- CI/CD
- Observability
- Secure development
- Proven ability to troubleshoot complex production issues through experimentation, profiling, and root-cause analysis.
- Experience with workflow orchestration, batch processing, and job scheduling frameworks.