Apply with hirly
Machine Learning Engineer, Predictive Maintenance
Assetwatch · United States
Upload your resume to see how well you match this job — free, in seconds, no account needed.
Your resume is used only to score it against this job. If you don't create an account, it is deleted within 24 hours.
Already have an account? Sign in to see your saved application
AssetWatch serves global manufacturers by powering manufacturing uptime through the delivery of an unparalleled condition monitoring experience, with a passion to care about the assets our customers care for every day. We are a devoted and capable team that includes world-renowned engineers and distinguished business leaders united by a common goal – To build the future of predictive maintenance. As we enter the next phase of rapid growth, we are seeking people to help lead the journey.
Job Summary
We are seeking an experienced Machine Learning Engineer with a specialized background in Signal Processing and Predictive Maintenance. The ideal candidate will have a talent for extracting diagnostic features and fault signatures from IIoT sensors, including accelerometers, temperature sensors, and electrical signals, as well as other static and contextual machine data such as images and equipment metadata. The successful candidate will utilize this data to develop and implement algorithms tailored for the diagnosis, prognosis, anomaly detection, and health assessment of critical machinery components, including bearings, gearboxes, shafts, motors, pumps, fans, belts, and more. This crucial role requires a thorough understanding of classical Machine Learning, Artificial Neural Networks, and Deep Learning algorithms, including CNN and Sequence-based Models such as RNNs and LSTMs enhanced with attention mechanisms. In addition, expertise in supervised and unsupervised learning, anomaly detection, ranking and decision models, feature engineering, and leveraging Large Language Models (LLMs) and agentic AI is a must. This is a hands-on role that also carries solution-architect responsibility: the candidate is expected to design end-to-end ML solutions — from data access and large-scale feature computation through training, inference, and production monitoring — on AWS, using services such as SageMaker (Training, Processing, Pipelines, Endpoints), S3, Athena, AWS Glue, EMR/Spark, Lambda, Step Functions, ECS/Fargate, Timestream, Aurora/RDS, Amazon Bedrock, and CloudWatch. Strong software engineering fundamentals and data/ML engineering at scale (distributed processing with Spark or equivalent) are also essential for building maintainable, testable, and production-ready ML systems. While mastery in algorithm and model development is a must, experience developing and deploying these solutions within the AWS cloud environment and collaborating closely with MLOps and Data Engineering teams is required.
Essential Responsibilities
- Extract, preprocess, and analyze data from various sensor sources, primarily machine vibration, while also working with temperature, electrical, image, and machine-context data to identify and enhance diagnostic features and fault-specific signatures.
- Determine the appropriate modeling techniques for each problem, including classical Machine Learning, supervised and unsupervised learning, anomaly detection, Artificial Neural Networks, Deep Learning, signal-processing methods, and hybrid or rules-based approaches.
- Train, validate, and deploy predictive maintenance models to accurately identify and predict machinery faults such as bearing, gearbox, shaft, motor, pump, fan, and belt faults. This may include CNNs, RNNs, LSTMs, attention-based models, and other time-series or representation-learning approaches.
- Develop and evaluate time-domain, frequency-domain, and contextual features, including harmonics, sidebands, envelope-spectrum features, running-speed characteristics, bearing and gear-mesh fault frequencies, and other fault-specific evidence.
- Act as the solution architect for assigned ML initiatives: define the end-to-end design — data sources and contracts, batch vs. near-real-time execution, feature computation strategy, storage layout (S3 partitioning, Timestream, Aurora/MySQL), orchestration, inference pattern, and cost/scalability tradeoffs — and document the design before large-scale execution, since backfills across tens of thousands of assets are expensive to rerun.
- Build and scale distributed data and feature pipelines over large historical sensor datasets using Spark (EMR, Glue, or SageMaker Processing), Athena, and Parquet-based data layouts, with attention to partitioning, throughput, and compute cost.
- Design and run data mining and annotation workflows that connect Condition Monitoring Engineers with Data Science — including case discovery, labeling interfaces and conventions, fault-type and time-window labeling, label QA, and turning expert feedback into reusable supervised training data.
- Build and maintain benchmark, training, validation, and regression datasets using real machine cases; work with Condition Monitoring Engineers and other domain experts to establish reliable ground truth and measurable model acceptance criteria.
- Design experiments and compare new algorithms against existing production methods, investigating false positives, false negatives, suppression behavior, data-quality issues, and performance across different machines, operating conditions, sampling configurations, and sensor characteristics.
- Develop robust, modular, reusable, and maintainable ML software in Python, including production-quality model code, shared libraries, configuration, unit and regression tests, code reviews, refactoring, debugging, and version-controlled development workflows.
- Use modern AI-assisted and agentic coding approaches to accelerate data preparation, feature engineering, model prototyping, testing, debugging, pipeline development, and documentation; develop or integrate LLM- and agent-based decision-support capabilities where they provide measurable value.
- Collaborate with the research and engineering team to constantly refine and improve model architectures, algorithms, and signal-processing approaches, ensuring high accuracy, explainability, scalability, and maintainability.
- Work closely with MLOps and Data Engineering teams to ensure smooth deployment of models, feature pipelines, data contracts, and other ML solutions in production environments, including CI/CD for ML, model versioning and registry, IaC (CloudFormation/CDK/Terraform), containerization (Docker/ECR), and production monitoring and alerting via CloudWatch, and participate in production validation and troubleshooting.
- Stay updated with the latest advancements in Machine Learning, Deep Learning, Signal Processing, agentic AI, and industrial condition monitoring, ensuring our solutions remain at the forefront of the industry.
- Present findings, model performance, strategies, tradeoffs, and solutions to other teams and stakeholders in a clear and concise manner, ensuring that insights drive actionable outcomes. REQUIREMENTS
- Master's or Ph.D. in Mechanical Engineering, Electrical Engineering, Computer Science/Engineering, or a related field.
- Proven experience in Predictive Maintenance and Condition Monitoring with a focus on Signal Processing techniques such as FFT, Short Time Fourier Transform, Time Synchronous Averaging, Wavelet Transform, Spectral Kurtosis, Spectral Correlation, Envelope Analysis, Hilbert Transform, and related methods.
- Strong foundations in Machine Learning and model development, including classical Machine Learning algorithms, supervised and unsupervised learning, anomaly detection, feature engineering, Artificial Neural Networks, Convolutional Neural Networks, Sequence-based Models such as RNNs and LSTMs, and Attention Mechanisms.
- Strong understanding of time-series modeling and evaluation, including the ability to select appropriate techniques based on the physical problem, data characteristics, available ground truth, and operational requirements.
- Hands-on experience architecting and delivering production ML solutions on AWS is required, including several of: SageMaker (Training, Processing, Pipelines, Model Registry, Endpoints), S3, Athena, AWS Glue, EMR, Lambda, Step Fun