Data/ML Engineer

Requisition # 2026-22319
Date Posted 3 hours ago(10/7/2026 3:12 PM)
Department
Schl of Public & Int'l Affairs
Category
Information Technology
Job Type
Full-Time

Overview

The Accelerator seeks a Data/ML Engineer to strengthen our data team and advance the engineering, enrichment, and provisioning of the data we collect.

 

The Accelerator at Princeton includes a portfolio of multiple planned independent and intersecting tools, built on a shared data and compute platform serving computational social scientists at research institutions across North America, Europe, and Africa. The Data/ML Engineer will work within our team to help drive data engineering and machine learning initiatives and collaborations. They will play a crucial role in building and operating the pipelines that transform large-scale social media and web behavior data into research-ready data products, and in developing the machine learning and enrichment capabilities that extend their value. They will work on problems that have no precedent and little source material, requiring novel solutions. They will also be responsible for working with the other teams within the Accelerator and our external partners to help foster collaboration and create an impactful environment for our users.

 

This is a 6-month term role with potential for extension. A remote work arrangement within the United States may be considered for candidates with the appropriate background and experience.

Responsibilities

Strategy:

  • Work closely with the Accelerator leadership team to align data engineering and machine learning initiatives with overarching goals and long-term vision.
  • Identify and prioritize development projects that benefit from data engineering and machine learning methodologies and innovations.

 

Data Engineering:

  • Design, build, and operate data pipelines across the Accelerator's medallion architecture, with end-to-end ownership of transformation layers that serve researchers.
  • Ensure the accuracy, integrity, and quality of data to be made available through the Accelerator, including data quality validation at pipeline boundaries and enforcement of versioned schema contracts.
  • Diagnose and optimize distributed data processing workloads at production scale.
  • Develop deployment automation, CI/CD, and release processes for data products, including versioned data releases and researcher-facing change documentation.

 

Machine Learning and Data Science:

  • Design, develop, and operate ML and NLP enrichment pipelines over large-scale text and behavioral data, including language identification, translation, and topic and content classification.
  • Own the full lifecycle of enrichment models: selection, evaluation against labeled data, batch inference architecture, cost efficiency, and reprocessing and versioning strategy.
  • Develop ML-ready feature layers and data products to support advanced research use cases.
  • Evaluate and apply large language model workflows and other emerging AI methods where they demonstrably improve outcomes, with attention to their validity for downstream scientific analysis.
  • Apply statistical analysis and modeling to characterize datasets, estimate coverage, and support research design.

 

Platform Operations and Cost Engineering:

  • Contribute to cost attribution, visibility, and governance across institutional workspaces, including cluster policies, budget controls, and storage lifecycle management.
  • Design data and ML workloads to operate within the platform's cost governance framework.
  • Develop automation for workspace and project provisioning as institutions and research projects onboard.
  • Operate within Unity Catalog governance, multi-tenant isolation, and research data security requirements.

 

Research and Collaboration:

  • Work effectively in a modern, professional software and data engineering environment with a strong understanding of Agile concepts and practices.
  • Modern Software Engineering Foundations: agile (Scrum), DevOps, CI/CD, code review, and pair programming, with working knowledge of cloud compute platforms to support collaborative, scalable, and efficient development.
  • Author and maintain researcher-facing documentation and provide direct technical support to research users of the platform.
  • Collaborate with research teams to define data products, sampling frames, and enrichment requirements, and apply state-of-the-art techniques to ongoing scientific challenges.
  • Stay current with the latest advancements in data engineering, machine learning, and relevant fields to continuously innovate.
  • Build strong relationships with external partners, driving collaborations that enhance the Accelerator's scientific impact.

Qualifications

Essential Qualifications:

  • 3+ years of relevant experience as a data engineer, machine learning engineer, or data scientist, which may include graduate research and internship experience, with a record of building production systems that operate reliably at scale. Experience working in a remote, agile environment.
  • Bachelor's degree or equivalent in a relevant field.
  • Strong proficiency in Python and SQL, and hands-on experience with distributed data processing (e.g., Apache Spark) on large data volumes.
  • Experience building, evaluating, and operating machine learning or NLP pipelines, including batch inference.
  • Working knowledge of cloud data platforms.
  • Strong communication and interpersonal skills to effectively collaborate with researchers in the field, other engineers at various levels of experience, and administrative and leadership team members.

 

Preferred Qualifications:

  • Experience with Azure and Databricks, including Unity Catalog.
  • Experience with infrastructure-as-code (e.g., Terraform), containers, and CI/CD tooling.
  • Experience with large-scale social media, web behavior, or text-as-data research.
  • Familiarity with large language model annotation workflows and their evaluation.
  • Publications in reputable scientific journals or conferences is desirable.

 

Princeton University is an Equal Opportunity Employer and all qualified applicants will receive consideration for employment without regard to age, race, color, religion, sex, sexual orientation, gender identity or expression, national origin, disability status, protected veteran status, or any other characteristic protected by law.

 

The University considers factors such as (but not limited to) scope and responsibilities of the position, candidate's qualifications, work experience, education/training, key skills, market, collective bargaining agreements as applicable, and organizational considerations when extending an offer. The posted salary range represents the University's good faith and reasonable estimate for a full-time position; salaries for part-time positions are pro-rated accordingly.

 

If the salary range on the posted position shows an hourly rate, this is the baseline; the actual hourly rate may be higher, depending on the position and factors listed above.

 

The University also offers a comprehensive benefit program to eligible employees. Please see this link for more information.

Standard Weekly Hours

36.25

Eligible for Overtime

No

Benefits Eligible

Yes

Probationary Period

180 days

Essential Services Personnel (see policy for detail)

No

Physical Capacity Exam Required

No

Valid Driver’s License Required

No

Experience Level

Mid-Senior Level

#Ll-DP1

Salary Range

$120,000 to $135,000

Options

Sorry the Share function is not working properly at this moment. Please refresh the page and try again later.
Share on your newsfeed

Connect With Us!

Join our Talent Network to receive updates about working at Princeton.

Princeton University job offers are contingent upon the candidate’s successful completion of a background check, reference checks, and pre-employment screening, as applicable.


If you have questions or comments regarding the iCIMS Privacy Policy or iCIMS FAQs, please contact accounts@icims.com.


Go to our careers site.