EB-2 NIW Data Engineer

What Cancer Detection and Wildfire Prevention Have in Common: The Person Who Cleans and Connects the Data

The Work That Powers Machine Learning

Machine learning models make predictions. But before any model can make a prediction, someone has to build the pipeline that collects the data, cleans it, transforms it from its raw form into a format the model can use, and keeps it flowing reliably at scale. That work is data engineering. It is not glamorous. Most people using AI systems never think about it. Without it, machine learning does not work, which makes this a strong EB-2 NIW data engineer foundation.

He has been doing that work for 21 years. Not in one domain - across banking, financial services, media, e-learning, and now sports analytics. The domain changes. The underlying discipline does not. His current role at a sports data analytics company involves preparing and maintaining machine-learning training datasets for sports performance prediction. He built a new Data Warehouse on Google BigQuery for the R&D team that reduced ML data preparation time by 70%, and optimized the table architecture to cut cloud processing costs by 50%. He is currently the only data engineer in the R&D department where all outcomes of future games and matches are predicted, strengthening the EB-2 NIW data engineer case.

His proposed endeavor takes that same data pipeline discipline and applies it to two problems where the data integration challenge is the limiting factor: early cancer detection and forest fire prevention. The machine learning techniques for early cancer detection exist. Multi-source sensors for wildfire monitoring exist. What does not yet exist, in a production grade, reliable, scalable form, is the data infrastructure that makes these models viable in real-world operational settings. That is what his 21 years of work has been building as an EB-2 NIW data engineer.

Two Systems, One Skill Set

The Cancer Detection and Prevention System and the Forest Fire Alert and Prevention System look very different on the surface for an EB-2 NIW data engineer. One draws from clinical records, genomics, lifestyle factors, and wearable device streams. The other draws from drone thermal cameras, LiDAR, multispectral sensors, satellite imagery, and weather feeds. Different domains, different data types, different end users.

The underlying engineering problems are the same.

Both require multi source data integration - combining data from incompatible systems with different schemas, update frequencies, and quality levels into a unified format. Both require real-time processing pipelines wearable health monitors and drone sensors both produce continuous data streams that cannot wait for batch processing. Both require ML ready data preparation - cleaning, normalizing, feature engineering, and quality validation before model training or inference. Both require scalable architecture that can handle growing data volumes without rebuilding the pipeline each time, strengthening the EB-2 NIW data engineer argument.

His entire career is the preparation for both. The specific applications are new. The engineering methodology is not for an EB-2 NIW data engineer.

Twenty-One Years Building the Same Skill at Increasing Scale

He started in 2003 as an IT Software Developer at a manufacturing company in Vietnam, monitoring an ERP system and building the first reporting infrastructure on top of it. Two years later he moved into database development, joining a U.S. headquartered e-learning company’s Vietnam operations where he designed OLTP databases, data warehouses, OLAP cubes, and ETL processes, and implemented database mirroring for high availability. From there, four years at a U.S. headquartered data services company’s Vietnam branch building banking client database architectures designed to handle up to one billion records per month with slowly changing dimension tracking.

He then spent four years as Senior Database Administrator at a UK-headquartered technology services company’s Vietnam branch, designing data warehouses for multi-timezone and multi-fiscal-year data, resolving live database issues with profiler analysis, and building migration tools to consolidate four legacy systems into one. He managed the team on Scrum/Agile and implemented MicroStrategy BI for executive reporting. Two years as Lead Data Consultant at a U.S.-headquartered company’s Vietnam operation added natural language processing, social media data extraction (Python/NLTK/Gensim), Hadoop/Hive big data workflows, and data visualization to his toolkit.

In Canada, he moved through a financial data solutions company (where he migrated 2 billion records within 8 hours for a system cutover), a data integration firm, and then into his current role at the sports analytics company where he does precisely the work his proposed endeavor requires: preparing raw historical data for machine learning models, optimizing query performance, and building the pipeline infrastructure that makes ML predictions reliable and scalable.

The National Importance Case for Both Systems

The U.S. cancer burden is documented and severe, making this directly relevant to an EB-2 NIW data engineer case. Approximately 1.9 million new cancer cases are diagnosed annually. The five-year survival rate for localized cancer is dramatically higher than for distant-stage cancer - in some cancers, the difference between early and late detection is the difference between 100% and 32% survival. The national cancer care expenditure exceeded $208 billion in 2020. Multiple federal programs - the Cancer Moonshot initiative, the National Cancer Plan, ARPA-H, NIH’s 2021-2025 strategic plan - explicitly identify AI-enabled early detection as a national priority.

The wildfire burden is equally documented. The 2020 wildfire season burned over 10.3 million acres, destroyed nearly 10,000 buildings, and caused approximately $16.5 billion in economic losses. The USDA Forest Service’s Wildfire Crisis Strategy, FEMA’s disaster management priorities, Executive Order 14008 on climate resilience, and the $3.5 billion in the Infrastructure Investment and Jobs Act for wildfire risk reduction all name early detection technology including drones, sensors, AI, and satellite integration as the priority intervention. The DOI has expanded satellite-based wildfire detection as part of the Biden-Harris administration’s Investing in America agenda, strengthening the EB-2 NIW data engineer national-importance argument.

His proposed systems address both directly. The Cancer Detection System integrates the multi-source clinical data the federal programs are funding research to improve. The Forest Fire Alert System operationalizes the sensor and ML approaches that federal wildfire strategy has identified as the priority tool. The data engineering methodology is his contribution as an EB-2 NIW data engineer: the infrastructure layer that makes these systems production-viable rather than research prototypes.

The Quantified Career Achievements

His career record includes specific, measurable accomplishments at each stage for an EB-2 NIW data engineer.

20,000 billed hours delivered on ETL system project in lead and developer position, supporting the EB-2 NIW data engineer profile.

70% reduction in ML data preparation effort for the R&D team by building a new Google BigQuery Data Warehouse that consolidates data from multiple disparate sources into a clean, ML-ready format.

90% SQL query performance improvement by identifying and resolving a critical tempdb capacity constraint affecting large transaction processing.

50% reduction in Google Big Query processing costs by optimizing table schema, partition strategies, and clustered key design.

30% improvement in SSIS importer performance for financial data file loading.

40% reduction in client effort for month-end reporting by delivering a purpose-built reporting tool.

2 billion records migrated within 8 hours to meet a hard deadline for financial system cutover, with zero data integrity failures, strengthening the EB-2 NIW data engineer case.

One billion records per month database architecture designed and implemented for banking clients at a U.S.-headquartered data services company.

How the Petition Was Built

This was a direct petition. The career record, quantified accomplishments, and technical credentials were already in place.

- National importance sourcing: American Cancer Society (1.9M annual cases, $208B expenditure), CDC survival rate data, Cancer Moonshot/NCI National Cancer Plan, ARPA-H establishment, NIH 2021-2025 strategic plan, NCI budget ($2.5B, +$500M), Executive Order 14091 on health equity; NFPA/NIFC wildfire data (7-10M+ acres annually), DOE/FEMA wildfire priorities, USDA Forest Service Wildfire Crisis Strategy, Executive Order 14008, IIJA $3.5B wildfire allocation, DOI satellite expansion initiative, National Cohesive Wildland Fire Management Strategy.

- Well-positioned evidence: 21 years progressive data engineering, current ML data pipeline at sports analytics company, 70%/90%/50% quantified performance improvements, multi-domain experience (banking, financial, media, e-learning, sports analytics, environmental), expertise in Google BigQuery, Apache Hadoop/Hive, SSIS, ML data preparation, Python, R.

- Proposed system technical frameworks: Apache Kafka, Spark Streaming, BigQuery, S3 for cancer system; CNN/RNN hybrid modeling, multispectral/thermal drone data, real-time monitoring with Apache Kafka/Flink for wildfire system.

I-140 filed as a self-petition without a U.S. employer.

The Outcome

Approved.

A self-petitioned EB-2 NIW for a data engineer who spent 21 years building the data pipelines that make machine learning work from one-billion record banking architectures in Vietnam to production ML systems in Canada proposing to apply that same infrastructure discipline to early cancer detection and wildfire prevention, two of the most directly federally prioritized AI applications in the United States.

For Data Engineers and ML Infrastructure Professionals

If your career is in data engineering, ETL architecture, data warehouse design, or ML data preparation and you are proposing to apply those skills to documented U.S. national priorities in healthcare AI, environmental data systems, or related domains, the NIW is worth a serious assessment. The Dhanasar test evaluates the national importance of the proposed endeavor and your positioning to advance it. A 21 year career building production-grade data infrastructure for ML systems, with quantified performance outcomes across multiple industries, is a strong answer to both.

Questions Data Engineers Ask Us

Can a data engineer or ETL specialist qualify for an EB-2 NIW?

Yes. The EB-2 NIW evaluates whether the proposed endeavor has substantial merit
and national importance, and whether the petitioner is positioned to advance it. A proposed endeavor that applies data engineering and ML data pipeline expertise to federally-prioritized applications, early cancer detection (Cancer Moonshot, NCI, ARPA-H) and wildfire prevention (USDA Wildfire Crisis Strategy, FEMA, IIJA), satisfies the national importance prong. A 21-year career with documented, quantified performance improvements in production ML data systems satisfies the well-positioned prong.

How does data engineering expertise connect to cancer detection or wildfire prevention specifically?

The limiting factor in deploying ML-based healthcare and environmental monitoring systems is not the modeling algorithm, it is the data pipeline. Early cancer detection requires integrating clinical records, genomics, wearable device streams, lifestyle factors, and environmental indicators from incompatible sources into a format that ML models can use reliably. Wildfire alert systems require fusing drone thermal data, multispectral imagery, satellite feeds, and weather data in real time. Both require exactly the skills data engineers develop: multi-source integration, real-time processing architecture, data quality management, and ML-ready transformation pipelines. A data engineer who has built billion-record production systems across multiple industries is positioned to build the infrastructure layer that makes these applications viable at scale.

Does a career spanning Vietnam, Canada, and U.S.-headquartered employers transfer to a U.S. NIW case?

Yes, particularly when the career shows progressive seniority and the proposed U.S. endeavor draws directly from the accumulated skills. Multiple prior employers were U.S.-headquartered companies with Vietnam operations (California, Dallas, Seattle headquarters), which provides direct exposure to U.S. technology standards and business practices. The 21-year progression from entry-level software development to sole R&D data engineer for a live ML production system is the track record USCIS evaluates. The geography where that work was performed is less relevant than what the work demonstrated.

How does a two-track proposed endeavor (cancer detection plus wildfire prevention) affect the NIW case?

A dual-track proposed endeavor can strengthen the national importance case by demonstrating alignment with multiple documented federal priorities, provided both tracks share a coherent technical foundation. In this case, both systems require the same underlying data engineering methodology: multi-source integration, real-time streaming pipelines, ML-ready data preparation, and scalable cloud architecture. That unifying foundation means the petitioner is not claiming expertise in two completely different fields; they are proposing to apply one technical discipline (data engineering for ML) to two nationally important domains. The petitioner must demonstrate well-positioning for both tracks, which a career spanning banking, environmental, sports analytics, and ML data preparation - all requiring the same core pipeline skills - does.

Does work history of building ML data pipelines for non-healthcare or non-environmental applications help an NIW for healthcare and environmental AI systems?

Yes. USCIS evaluates whether the petitioner can advance the proposed endeavor, not whether they have done so in the identical application domain. A data engineer who has built production ML training data pipelines with documented 70% efficiency improvements, 50% cost reductions, and sole responsibility for a live prediction system has demonstrated the specific technical capability (ML data pipeline at scale) that the proposed cancer detection and wildfire systems require. The transferability of data engineering skills across domains is well-established, and the track record of quantified performance outcomes in ML data preparation is directly relevant evidence.

If your career is in data engineering and you want to understand whether your background and proposed work support an EB-2 NIW, start with a free assessment.
Free assessment: immignis

Don't guess your eligibility. Get a free, expert assessment today.

You may qualify and not even know it yet.

Submit Your Free Assessment Request