Rawpixel 653764 unsplash — Vassar Labs

Job Description

We are looking for a skilled Python Developer with 3+ years of work experience to design, develop, and maintain scalable applications and backend services. The ideal candidate should have strong Python programming skills, good problem-solving ability, and experience working with APIs, databases, and modern development frameworks.

Key Responsibilities:
  • Design, build, and operate distributed batch and streaming pipelines that process highvolume data.
  • Build high-performance, production-grade REST (and where appropriate, gRPC) APIs that serve data to web, mobile, and AI applications.
  • Work with Kafka or rabbitmq like similar message queues, Spark/PySpark, and Flink/Storm to build ingestion, enrichment, and aggregation workflows with strong guarantees around ordering, idempotency, and fault tolerance.
  • Model and optimise data in RDBMS database and NoSQL database and object storage
    (MinIO/S3), including partitioning, indexing, and spatial query tuning.
  • Profile and optimise Python code for throughput and memory: async I/O, multiprocessing, vectorisation, and knowing when to drop to a faster tool.
Preferred Skills :
  • Professional software development, with Python as your primary language.
  • Deep Python expertise: concurrency (asyncio, threading, multiprocessing, the GIL), typing, packaging, testing, and performance profiling.
  • Proven experience building and running distributed data pipelines in production, with at least one of Spark, Flink, Kafka Streams, Beam, Storm or similar.
  • Hands-on experience with Kafka or another message broker, including consumer groups, partitioning, delivery semantics, and back-pressure.
  • Strong API design skills using FastAPI, Django/DRF, or Flask: versioning, pagination, auth
    (OAuth2/JWT), rate limiting, and OpenAPI documentation.
  • Solid SQL and PostgreSQL knowledge, including query plans, indexing, and schema design.
  • A strong grasp of distributed systems fundamentals: consistency, partitioning, retries,
    exactly-once vs at-least-once, and failure modes
  • Time-series and IoT data handling (TimescaleDB, InfluxDB, or high-frequency telemetry).
  • Workflow orchestration with Airflow, Prefect, or Dagster.
  • Exposure to ML pipelines or serving models behind APIs

Get In Touch