5.6 KiB
Market evidence report — data-scientist
Source: 46 real job ads (JSearch API, countries: us 46), extracted into the MSSQL evidence store; as of 2026-07-11. This report contains extracted, aggregated facts only — no ad text is reproduced (copyright / platform terms).
Seniority distribution
| Seniority | Ads | Share |
|---|---|---|
| mid | 25 | 54 % |
| lead | 10 | 22 % |
| senior | 8 | 17 % |
| junior | 3 | 7 % |
Tools — full market ranking
| # | Item | Ads | Share |
|---|---|---|---|
| 1 | Python | 36 | 78 % |
| 2 | SQL | 21 | 46 % |
| 3 | Spark | 11 | 24 % |
| 4 | AWS | 10 | 22 % |
| 5 | EMR | 5 | 11 % |
| 6 | Hadoop | 5 | 11 % |
| 7 | scikit-learn | 5 | 11 % |
| 8 | Azure | 4 | 9 % |
| 9 | Conda | 4 | 9 % |
| 10 | H2O | 4 | 9 % |
| 11 | Hive | 4 | 9 % |
| 12 | Kafka | 4 | 9 % |
| 13 | PyTorch | 4 | 9 % |
| 14 | Tableau | 4 | 9 % |
| 15 | TensorFlow | 4 | 9 % |
| 16 | Java | 3 | 7 % |
| 17 | MATLAB | 3 | 7 % |
| 18 | NoSQL | 3 | 7 % |
| 19 | pandas | 3 | 7 % |
| 20 | Power BI | 3 | 7 % |
| 21 | R programming language | 3 | 7 % |
Hard skills — full market ranking
| # | Item | Ads | Share |
|---|---|---|---|
| 1 | machine learning | 31 | 67 % |
| 2 | data analysis | 26 | 57 % |
| 3 | statistical modeling | 17 | 37 % |
| 4 | statistical analysis | 15 | 33 % |
| 5 | predictive modeling | 13 | 28 % |
| 6 | data mining | 12 | 26 % |
| 7 | data visualization | 9 | 20 % |
| 8 | natural language processing | 9 | 20 % |
| 9 | data modeling | 8 | 17 % |
| 10 | deep learning | 8 | 17 % |
| 11 | quantitative analysis | 8 | 17 % |
| 12 | time series analysis | 7 | 15 % |
| 13 | feature engineering | 6 | 13 % |
| 14 | anomaly detection | 5 | 11 % |
| 15 | classification | 5 | 11 % |
| 16 | clustering | 5 | 11 % |
| 17 | data cleaning | 5 | 11 % |
| 18 | data exploration | 5 | 11 % |
| 19 | model validation | 5 | 11 % |
| 20 | sentiment analysis | 5 | 11 % |
| 21 | algorithm development | 4 | 9 % |
| 22 | data integration | 4 | 9 % |
| 23 | data science | 4 | 9 % |
| 24 | risk assessment | 4 | 9 % |
| 25 | backtesting | 3 | 7 % |
| 26 | data analytics | 3 | 7 % |
| 27 | data cleansing | 3 | 7 % |
| 28 | data engineering | 3 | 7 % |
| 29 | data management | 3 | 7 % |
| 30 | data normalization | 3 | 7 % |
| 31 | predictive analytics | 3 | 7 % |
Methods — full market ranking
| # | Item | Ads | Share |
|---|---|---|---|
| 1 | machine learning | 4 | 9 % |
| 2 | backtesting | 3 | 7 % |
| 3 | mlops | 3 | 7 % |
Responsibilities — full market ranking
| # | Item | Ads | Share |
|---|---|---|---|
| 1 | data analysis | 16 | 35 % |
| 2 | model development | 11 | 24 % |
| 3 | solution development | 5 | 11 % |
| 4 | anomaly detection | 4 | 9 % |
| 5 | data retrieval | 4 | 9 % |
| 6 | client communication | 3 | 7 % |
| 7 | data combination | 3 | 7 % |
| 8 | data exploration | 3 | 7 % |
| 9 | data visualization | 3 | 7 % |
| 10 | documentation maintenance | 3 | 7 % |
| 11 | insight revelation | 3 | 7 % |
| 12 | model validation | 3 | 7 % |
| 13 | solution deployment | 3 | 7 % |
| 14 | trend identification | 3 | 7 % |
Regional breakdown
Corpus note: 46 relevant ads in total — below the 100-ad target for a fully reliable ranking. Percentages above should be read as indicative.
US (us)
46 ads.
Top hard skills:
- machine learning — 67 % (31 ads)
- data analysis — 57 % (26 ads)
- statistical modeling — 37 % (17 ads)
- statistical analysis — 33 % (15 ads)
- predictive modeling — 28 % (13 ads)
- data mining — 26 % (12 ads)
- data visualization — 20 % (9 ads)
- natural language processing — 20 % (9 ads)
- data modeling — 17 % (8 ads)
- deep learning — 17 % (8 ads)
Top tools:
- Python — 78 % (36 ads)
- SQL — 46 % (21 ads)
- Spark — 24 % (11 ads)
- AWS — 22 % (10 ads)
- EMR — 11 % (5 ads)
- Hadoop — 11 % (5 ads)
- scikit-learn — 11 % (5 ads)
- Azure — 9 % (4 ads)
- Conda — 9 % (4 ads)
- H2O — 9 % (4 ads)
Seniority: mid 54 % · lead 22 % · senior 17 % · junior 7 %
UK (gb)
Insufficient evidence — 0 ads (minimum for a regional ranking: 30). No ranking is reported for this region.
EU/DACH (de, at, ch, nl)
Insufficient evidence — 0 ads (minimum for a regional ranking: 30). No ranking is reported for this region.
Job title variants in the market
| Title | Ads |
|---|---|
| Data Scientist | 2 |
| Data Scientist - Experienced to Expert Level (Maryland) | 2 |
| Data Scientist, AWS Security | 2 |
| Data Scientist, Mid | 2 |
| Senior Associate, Data Scientist - Corporate Strategy | 2 |
| AI and ML Data Scientist Jobs | 1 |
| Data Science Lead, Personnel Health Research & Data Analytics | 1 |
| Data Scientist - Advanced Manufacturing Technologies (AMT) | 1 |
| Data Scientist - Expert | 1 |
| Data Scientist - Mid | 1 |
| Data Scientist - Senior | 1 |
| Data Scientist - Sr | 1 |
| Data Scientist - Supply Chain Risk Management with Security Clearance | 1 |
| Data Scientist - Traditional and Generative AI | 1 |
| Data Scientist (U.S. Citizen, Secret/Top Secret), Washington D.C. | 1 |
| Data Scientist II | 1 |
| Data Scientist III | 1 |
| Data Scientist Jobs | 1 |
| Data Scientist Lead ID | 1 |
| Data Scientist SETA (TS/SCI #26-099) | 1 |
| Data Scientist, Level 4 with Security Clearance | 1 |
| Data Scientist, Mid level | 1 |
| Data Scientist, Product Analytics | 1 |
| Experienced Data Scientist | 1 |
| Junior DoD Data Scientist - Onsite in Alexandria | 1 |
Methodology: entities extracted per ad ({hard_skills, tools, methods, responsibilities, seniority}), normalized, counted as DISTINCT ads per entity; report threshold ≥ 3 ads. Headline sections in skills.md/tools.md use the stricter ≥ 20 % threshold.