Job Details

Job Title
Senior Data Engineer (Python)
Location
Noida
Job Description

JOB PROFILE

 

Position

Senior Python Data Engineer (Pipeline Optimization)

Location

Gurgaon

Reports to

 

Category

Data Engineering / Analytics

Reporting / Proficiency Level

NA

Level

Senior (5+ years)

Our Purpose

At Niva Bupa, our purpose is “to give every Indian the confidence to access the best healthcare” by empowering them with knowledge, guiding them with expertise, and providing them with a gamut of services that instils confidence and puts control back in their hands- just the way they want every moment of their life to be.

Our Values

  1. Commitment
  2. Innovation
  3. Empathy
  4. Collaboration
  5. Transparency

About Niva Bupa Health Insurance Company

 

Niva Bupa Health Insurance Company Limited is a Public Listed Company on Stock exchange(s). The company’s purpose is to give every Indian the confidence to access the best healthcare. It intends to play the role of an enabler in the lives of its customers and help them live life without constraints. This is reflected in its brand philosophy – ‘Zindagi Ko Claim Kar Le’.

 

As of December 31, 2025, Niva Bupa had over 210 physical branches across India. It additionally offers health insurance through its ecosystem partners including 2.2 Lakh agents, 575+ brokers, and over 115 Banca & Other Corporate Agency Partners. The company currently covers 24.5 million lives and has 10,587 hospitals empanelled in its hospital network.

 

Niva Bupa has consistently maintained 90%+ claim settlement ratio over the last 4 financial years, having ended Q3FY26 with claim settlement ratio of 94.1%. With an employee base of over 10,100 people, the company is a certified Great Place to Work six times in a row.

 

 

 

Key Roles & Responsibilities

Job Summary:

We are looking for a Senior Python Data Engineer who can take ownership of our large-scale data pipelines and make them run faster, leaner and more reliably. The person in this role will spend most of their time working on real production code, optimizing it end-to-end and bringing down pipeline runtimes that today stretch into many hours.

 

This is primarily a Python role, not a distributed-systems role. Our pipelines run on high-memory single-node servers (500GB+ RAM, 16+ CPUs), so the focus is on writing efficient Python code, managing memory carefully, parallelizing where it makes sense, and getting the most out of modern Python data libraries. Existing pandas-heavy code needs to be re-engineered into faster equivalents, and new pipelines need to be built with performance in mind from day one. You will be working closely with the actuarial, MIS and reporting teams.

 

The right person will be hands-on, comfortable reading and rewriting other people’s code, and curious enough to figure out why one approach takes 27 hours and another takes 90 minutes. Output equivalence is non-negotiable – whatever we ship has to match the original numbers exactly, every single time.

Key Responsibilities:

 

  1. Own end-to-end optimization of existing Python pipelines – profile them, identify the slow steps, and rewrite them so the same job that takes 20+ hours today finishes in a fraction of that time.
  2. Re-engineer pandas-heavy code into faster, more memory-efficient equivalents. Most of our existing scripts are written in pandas; we are moving towards modern Python data libraries (Polars, DuckDB, PyArrow) where they make sense.
  3. Build new data pipelines for various data process and similar workloads – from source extraction all the way to the final reporting layer.
  4. Handle large-volume data extraction from Oracle and PostgreSQL using Python connectors. Tune fetch parameters (arraysize, prefetchrows, chunking, parallel reads) to keep ingestion fast and predictable.
  5. Work with parquet and other columnar formats – partitioning strategies, compression, projection pushdown – to keep intermediate storage and read times under control.
  6. Manage memory carefully on high-RAM single-node servers (500GB+). This includes chunking large joins, releasing memory between steps, avoiding OOMs on multi-hundred-million-row datasets, and using parallelism (ThreadPoolExecutor, multiprocessing) where it actually helps.
  7. Ensure 100% output equivalence between the old and new versions of every pipeline. Build comparison and validation scripts so that nothing slips through during a rewrite.
  8. Write advanced SQL for both data validation and PostgreSQL output tables – window functions, CTEs, joins on tens of millions of rows, indexing, and basic query tuning.
  9. Translate business logic from actuarial, MIS and reporting teams into clean, testable Python code – reinsurance calculations, earned premium, member-based reporting, portfolio cuts and so on.
  10. Support monthly production runs – reconfiguring parameters, monitoring runtimes, fixing data issues as they come up, and keeping the reporting calendar on track.
  11. Maintain proper version control, logging and elapsed-time reporting on every script so that production behaviour is observable and reproducible.

 

Key Requirements – Education & Certificates

  1. B.E. / B.Tech / M.Tech / MCA in Computer Science, Information Technology, or any related discipline. Equivalent degrees in Statistics, Mathematics or Engineering are also fine if backed by strong hands-on data engineering experience.
  2. Certifications in Python, SQL or any cloud data platform (AWS / Azure / GCP) are good to have but not mandatory.

Key Requirements - Experience & Skills (Must have)

Must Have

  1. 5+ years of hands-on Python experience, with a strong portion of that spent on data engineering / ETL work – not just scripting around dashboards.
  2. Strong working knowledge of pandas – merges, group-bys, multi-index handling, the usual NaN / null gotchas, and a clear sense of where pandas hits its limits. Experience migrating pandas code to faster equivalents is a big plus.
  3. Solid SQL skills – should be comfortable writing and reading complex queries (joins, CTEs, window functions, aggregations) on large tables.
  4. Hands-on experience pulling data from Oracle and PostgreSQL using Python connectors (cx_Oracle / oracledb / psycopg2 / asyncpg etc.). Should understand fetch tuning, chunking and connection management for large volumes.
  5. Comfort with parquet and columnar file formats – reading, writing, partitioning, compression options, and projection / predicate pushdown.
  6. Real-world experience with memory optimization on single-node servers – chunked processing, releasing memory between steps, debugging OOMs, and using parallelism (ThreadPoolExecutor / multiprocessing) sensibly.
  7. Good debugging instincts – should be able to take an existing, undocumented Python script and figure out what it’s doing and where it’s slow.
  8. Comfort working in a Jupyter notebook based environment for development, and a habit of writing clean, version-controlled, well-logged code.

 

Good to Have

  1. Exposure to Polars, DuckDB or PyArrow. We don’t expect everyone to know these – if you do, it’s a strong plus; if you don’t, willingness to pick them up quickly is enough.
  2. Working knowledge of PySpark / Apache Spark. Useful for context and for any future distributed work, but the day-to-day on this role is single-node Python.
  3. Prior experience in health insurance, BFSI or any other regulated industry. Understanding of premium, claims, reinsurance or member-level reporting is a definite advantage.
  4. Familiarity with workflow / scheduling tools like Airflow and with Git-based version control.
  5. Some exposure to one of the major cloud platforms (AWS / Azure / GCP) and to Excel-heavy reporting workflows used by business teams.

 

Apply to Job