*** This role is supporting NVIDIA ***
Education:
- Bachelor's degree in computer science or related engineering field, or equivalent experience.
Required Qualifications:
- 4+ years of professional experience in data engineering or a closely related field.
- Experience designing, building, and operating production ETL/ELT pipelines for large, complex datasets.
- Strong SQL and Python proficiency for data pipelines, automation, and production software.
- Familiarity with CI/CD, testing frameworks, and version control (e.g., Git).
- Experience with workflow orchestration (e.g., Airflow, Dagster, or Databricks Jobs/Lakeflow).
- Demonstrated experience designing and maintaining data models, data pipelines, and databases (relational and/or lakehouse/warehouse).
- Experience designing and operating high-volume logging or telemetry systems, with a focus on efficiency and monitoring.
- Experience building and maintaining dashboards, reports, and alerting, and keeping metric definitions accurate over time.
- Hands-on experience with Databricks, Spark, Delta Lake, or similar lakehouse/big-data platforms.
- Experience with a major cloud platform (AWS, Azure, or GCP) and large-scale storage systems.
- Proficiency using AI tools for coding and data tasks, paired with strong first-principles reasoning about design, coding, and datasets.
- Strong problem-solving skills, including triaging and resolving complex bugs that span multiple systems and identifying root causes.
- Strong written and verbal communication and coordination skills, with a track record of working across internal and external partners.
- Ability to work independently, manage your own work, and proactively provide status updates.
Preferred Qualifications:
- 8+ years of experience in data engineering or analytics engineering.
- Experience building data pipelines specifically for annotation projects running in annotation platforms (human-in-the-loop / labeling workflows).
- Direct experience with SuperAnnotate.
- Experience supporting LLM/VLM or other ML model training data pipelines.
- Experience working with large-scale, high-volume, multi-modal datasets (e.g., text, image, video, audio), with attention to storage and processing efficiency.
- Front-end or full-stack experience (JavaScript/TypeScript/HTML) for internal tooling and lightweight data/annotation UIs.
As a data engineer, you will design, build, operate, and maintain the data infrastructure behind dataset creation and annotation workflows for GenAI models, working with research, product, and engineering teams to turn raw source data into structured datasets for analysis and model training.
The core of the job is data modeling, ETL, and the ongoing maintenance of pipelines, databases, dashboards, reports, and alerts. You will own these systems over their full lifecycle: getting requirements, designing and building them, keeping them running, monitoring them, and fixing them when they break. You will also write production software, support annotation data workflows, and take on other engineering and operational tasks as needed.
This role requires strong communication and coordination. You will work with multiple internal and external partners and are expected to make pipelines and automations easy to use, easy to understand, and reliable, and to document them accordingly.
We expect proficiency with AI tools for coding and data tasks, along with strong first-principles understanding of design, coding, and datasets. You should use AI tools to work efficiently and be able to verify and explain their output.
Responsibilities:
- Build and maintain data pipelines: Design, implement, and maintain pipelines and data-processing frameworks for dataset creation, annotation workflows, and large-scale data transfer and transforms across batch and incremental patterns. Keep them running, monitored, and reliable over time.
- Handle large-scale multi-modal datasets: Build and maintain pipelines and data models for large-scale multi-modal datasets (e.g., text, image, video, audio) at high volume, accounting for storage, throughput, and processing efficiency.
- Manage high-volume logging: Design and maintain high-volume logging and telemetry, where efficiency and monitoring are primary concerns; ensure logs are captured, stored, and queried cost-effectively and are usable for debugging and operational visibility.
- Maintain dashboards, reports, and alerts: Build and maintain metrics, dashboards, reports, and alerting that track throughput, quality, cost, and program health. Keep metric definitions and outputs accurate as sources and requirements change.
- Own data modeling and databases: Design and maintain data models, schemas, and databases (layered/medallion architectures, dimensional models, warehouse/lakehouse tables) that serve both business reporting and model-training consumption.
- Operate on lakehouse platforms: Build, optimize, and maintain workloads on Databricks and similar systems—tables, transformations, orchestration, governance, and cost/performance.
- Ensure data quality and reliability: Implement data quality checks, validation, monitoring, and alerting; troubleshoot pipeline and data issues and fix root causes.
- Support annotation data workflows: Build and maintain pipelines that move data into and out of internal and third-party annotation platforms, including pre-processing, post-processing, validation, and delivery of annotated datasets.
- Support annotation project UIs and workflows: As needed, take on engineering tasks related to annotation project UIs and workflows, including development, testing, and QA.
- Write production software: Develop backend services, automation, and tooling in Python, applying testing, CI/CD, and version control.
- Triage and solve complex problems: Diagnose and resolve complex bugs that span multiple systems, identify root causes, and implement fixes.
- Communicate and coordinate: Work with internal and external partners to define requirements, and make pipelines and automations easy to use, easy to understand, and reliable. Document systems, definitions, and operational procedures.
- Work independently: Scope ambiguous requests, propose designs, and pick up cross-functional engineering and operational tasks as the team's needs change. Manage your own work and proactively provide status updates.
- Integrate across infrastructure: Connect pipelines and tooling with cloud platforms (AWS/Azure/GCP), storage systems, and annotation platforms, with attention to reliability and observability.
- **Only those lawfully authorized to work in the designated country associated with the position will be considered.**
- **Please note that all Position start dates and duration are estimates and may be reduced or lengthened based upon a client’s business needs and requirements.**
Rose International has been great to me. I thank everyone there for all of their hard work; it has not gone unnoticed.
Melody, Consultant
I have been very pleased with my experience with Rose International. Everyone that I encountered was very helpful and courteous.
Stephanie, Consultant
Rose is an assembly of people grounded in honesty, truth and dignity for all of its employees and contractors.
Samba, Consultant
It is a great pleasure being a part of the Rose International Team.
Toni, Consultant
Working for Rose International was the most pleasant assignment I have ever had. They were always on top of situations when necessary, and very helpful. I was very proud to be an employee of Rose International, and would recommend anyone to try to work with them.
Melvon, Consultant
EMPLOYEE COMMENTS