Rawpixel 653764 unsplash — Vassar Labs

Job Description

We are looking for an AI Engineer with 3+ years of experience and with strong foundations in Python/software engineering, SQL and data systems, statistics, and machine learning to design, develop, and deploy production-grade AI/ML solutions for real-world business and climate-tech applications. The role spans classical Machine Learning, Deep Learning, Generative AI, and production engineering, with an emphasis on correctness, scalability, maintainability, and measurable impact.

Key Responsibilities:
  • Design, develop, train, validate, and deploy AI/ML models for business and domain-specific use cases.
  • Write clean, efficient, testable Python and select appropriate synchronous, concurrent, asynchronous, or parallel execution patterns for data and inference workloads.
  • Work with structured and unstructured data including tabular data, text, images, geospatial data, and time-series data.
  • Build reliable data preparation workflows using Python and SQL, including data cleaning, transformation, feature engineering, and validation.
  • Apply sound ML methodology for model selection, training, hyperparameter tuning, cross-validation, error analysis, and evaluation.
  • Develop solutions using Machine Learning, Deep Learning, NLP, Computer Vision, and Generative AI/LLMs where appropriate.
  • Build and optimize AI pipelines, model inference workflows, REST APIs, and batch/stream processing components for production deployment.
  • Profile and optimize code, database queries, model inference, memory usage, throughput, and latency based on measurable bottlenecks.
  • Use software engineering practices such as modular design, unit/integration testing, version control, code review, logging, and reproducible experiments.
  • Work with cloud platforms, databases, containers, and MLOps tools to deploy and operate scalable AI solutions.
  • Monitor deployed models and services for model quality, data drift, reliability, scalability, latency, and resource utilization.
  • Collaborate with Data Scientists, Software Engineers, Product/Domain Experts, and Project Teams to translate requirements into robust AI solutions and maintain clear technical documentation.

Core Foundations & Required Skills:

Python & Software Engineering

  • Strong command of Python fundamentals including data structures, functions, OOP, modules/packages, exception handling, typing, iterators/generators, decorators, context managers, and the standard library.
  • Practical understanding of concurrency and parallelism: threading, multiprocessing, asyncio, concurrent.futures, synchronization primitives, queues, race conditions, deadlocks, and safe shared-state handling.
  • Understanding of the Python GIL and the ability to choose appropriate approaches for I/O-bound versus CPU-bound workloads.
  • Good knowledge of data structures, algorithms, time/space complexity, debugging, profiling, unit testing, and writing maintainable production code.
  • Comfort with Linux command-line workflows, Git-based development, REST APIs, and common software engineering practices.

SQL & Data Foundations

  • Strong SQL skills including joins, subqueries, CTEs, aggregations, GROUP BY/HAVING, window functions, conditional logic, and working with large datasets.
  • Understanding of relational database fundamentals including schema design, normalization, primary/foreign keys, transactions/ACID, indexes, and query execution plans.
  • Ability to diagnose and optimize slow queries and avoid common data-access problems such as unnecessary scans, repeated queries, and inefficient joins.
  • Hands-on data manipulation using libraries such as NumPy and Pandas, with awareness of vectorization, memory usage, missing data, outliers, and data quality checks.

Machine Learning & Statistics Foundations

  • Strong understanding of supervised and unsupervised learning, including regression, classification, clustering, dimensionality reduction, and common tree/ensemble methods.
  • Working knowledge of probability and statistics concepts used in ML, including distributions, sampling, descriptive statistics, correlation, hypothesis testing, and uncertainty.
  • Understanding of loss/objective functions, gradient-based optimization, bias-variance trade-off, overfitting/underfitting, regularization, feature selection, and hyperparameter tuning.
  • Strong model validation practices: train/validation/test splits, cross-validation, data leakage prevention, class imbalance handling, baselines, and reproducibility.
  • Ability to select and interpret appropriate evaluation metrics such as precision, recall, F1, ROC-AUC/PR-AUC, log loss, MAE/RMSE, and domain-specific metrics rather than relying on accuracy alone.

Deep Learning, AI & Production

  • Hands-on experience with Scikit-learn and at least one Deep Learning framework such as PyTorch or TensorFlow, with understanding of neural networks, backpropagation, optimizers, and training workflows.
  • Knowledge of one or more applied AI areas such as NLP, Computer Vision, time-series modelling, or geospatial ML; familiarity with modern architectures such as CNNs and Transformers is preferred.
  • Practical knowledge of Generative AI/LLMs, prompting, embeddings, retrieval, evaluation, and the limitations/risks of LLM-based systems.
  • Experience with model serving, APIs, Docker, cloud platforms, logging/monitoring, and basic MLOps practices for reliable production deployment.
  • Strong analytical problem-solving skills and the ability to explain technical trade-offs, debug failures systematically, and validate assumptions with data.

Good to Have:
  • Experience with LLM frameworks and tooling such as LangChain, LlamaIndex, Hugging Face, or equivalent.
  • Experience with RAG, vector databases, embeddings, AI agents, tool/function calling, and systematic LLM evaluation.
  • Exposure to distributed/data-processing systems such as Spark, Kafka, Ray, or equivalent.
  • Exposure to Azure/AWS/GCP AI and ML services, Kubernetes, CI/CD, and production observability.
  • Experience working with geospatial, satellite, climate, agriculture, water, or environmental datasets.
  • Knowledge of model/inference optimization techniques, GPU serving, batching, quantization, caching, and production-scale AI systems.

Get In Touch