# Market evidence report — data-scientist Source: **46 real job ads** (JSearch API, countries: us 46), extracted into the MSSQL evidence store; as of 2026-07-11. This report contains extracted, aggregated facts only — no ad text is reproduced (copyright / platform terms). ## Seniority distribution | Seniority | Ads | Share | |---|---|---| | mid | 25 | 54 % | | lead | 10 | 22 % | | senior | 8 | 17 % | | junior | 3 | 7 % | ## Tools — full market ranking | # | Item | Ads | Share | |---|---|---|---| | 1 | Python | 36 | 78 % | | 2 | SQL | 21 | 46 % | | 3 | Spark | 11 | 24 % | | 4 | AWS | 10 | 22 % | | 5 | EMR | 5 | 11 % | | 6 | Hadoop | 5 | 11 % | | 7 | scikit-learn | 5 | 11 % | | 8 | Azure | 4 | 9 % | | 9 | Conda | 4 | 9 % | | 10 | H2O | 4 | 9 % | | 11 | Hive | 4 | 9 % | | 12 | Kafka | 4 | 9 % | | 13 | PyTorch | 4 | 9 % | | 14 | Tableau | 4 | 9 % | | 15 | TensorFlow | 4 | 9 % | | 16 | Java | 3 | 7 % | | 17 | MATLAB | 3 | 7 % | | 18 | NoSQL | 3 | 7 % | | 19 | pandas | 3 | 7 % | | 20 | Power BI | 3 | 7 % | | 21 | R programming language | 3 | 7 % | ## Hard skills — full market ranking | # | Item | Ads | Share | |---|---|---|---| | 1 | machine learning | 31 | 67 % | | 2 | data analysis | 26 | 57 % | | 3 | statistical modeling | 17 | 37 % | | 4 | statistical analysis | 15 | 33 % | | 5 | predictive modeling | 13 | 28 % | | 6 | data mining | 12 | 26 % | | 7 | data visualization | 9 | 20 % | | 8 | natural language processing | 9 | 20 % | | 9 | data modeling | 8 | 17 % | | 10 | deep learning | 8 | 17 % | | 11 | quantitative analysis | 8 | 17 % | | 12 | time series analysis | 7 | 15 % | | 13 | feature engineering | 6 | 13 % | | 14 | anomaly detection | 5 | 11 % | | 15 | classification | 5 | 11 % | | 16 | clustering | 5 | 11 % | | 17 | data cleaning | 5 | 11 % | | 18 | data exploration | 5 | 11 % | | 19 | model validation | 5 | 11 % | | 20 | sentiment analysis | 5 | 11 % | | 21 | algorithm development | 4 | 9 % | | 22 | data integration | 4 | 9 % | | 23 | data science | 4 | 9 % | | 24 | risk assessment | 4 | 9 % | | 25 | backtesting | 3 | 7 % | | 26 | data analytics | 3 | 7 % | | 27 | data cleansing | 3 | 7 % | | 28 | data engineering | 3 | 7 % | | 29 | data management | 3 | 7 % | | 30 | data normalization | 3 | 7 % | | 31 | predictive analytics | 3 | 7 % | ## Methods — full market ranking | # | Item | Ads | Share | |---|---|---|---| | 1 | machine learning | 4 | 9 % | | 2 | backtesting | 3 | 7 % | | 3 | mlops | 3 | 7 % | ## Responsibilities — full market ranking | # | Item | Ads | Share | |---|---|---|---| | 1 | data analysis | 16 | 35 % | | 2 | model development | 11 | 24 % | | 3 | solution development | 5 | 11 % | | 4 | anomaly detection | 4 | 9 % | | 5 | data retrieval | 4 | 9 % | | 6 | client communication | 3 | 7 % | | 7 | data combination | 3 | 7 % | | 8 | data exploration | 3 | 7 % | | 9 | data visualization | 3 | 7 % | | 10 | documentation maintenance | 3 | 7 % | | 11 | insight revelation | 3 | 7 % | | 12 | model validation | 3 | 7 % | | 13 | solution deployment | 3 | 7 % | | 14 | trend identification | 3 | 7 % | ## Regional breakdown > **Corpus note:** 46 relevant ads in total — below the 100-ad target for a fully reliable ranking. Percentages above should be read as indicative. ### US (us) 46 ads. **Top hard skills:** - machine learning — 67 % (31 ads) - data analysis — 57 % (26 ads) - statistical modeling — 37 % (17 ads) - statistical analysis — 33 % (15 ads) - predictive modeling — 28 % (13 ads) - data mining — 26 % (12 ads) - data visualization — 20 % (9 ads) - natural language processing — 20 % (9 ads) - data modeling — 17 % (8 ads) - deep learning — 17 % (8 ads) **Top tools:** - Python — 78 % (36 ads) - SQL — 46 % (21 ads) - Spark — 24 % (11 ads) - AWS — 22 % (10 ads) - EMR — 11 % (5 ads) - Hadoop — 11 % (5 ads) - scikit-learn — 11 % (5 ads) - Azure — 9 % (4 ads) - Conda — 9 % (4 ads) - H2O — 9 % (4 ads) **Seniority:** mid 54 % · lead 22 % · senior 17 % · junior 7 % ### UK (gb) **Insufficient evidence** — 0 ads (minimum for a regional ranking: 30). No ranking is reported for this region. ### EU/DACH (de, at, ch, nl) **Insufficient evidence** — 0 ads (minimum for a regional ranking: 30). No ranking is reported for this region. ## Job title variants in the market | Title | Ads | |---|---| | Data Scientist | 2 | | Data Scientist - Experienced to Expert Level (Maryland) | 2 | | Data Scientist, AWS Security | 2 | | Data Scientist, Mid | 2 | | Senior Associate, Data Scientist - Corporate Strategy | 2 | | AI and ML Data Scientist Jobs | 1 | | Data Science Lead, Personnel Health Research & Data Analytics | 1 | | Data Scientist - Advanced Manufacturing Technologies (AMT) | 1 | | Data Scientist - Expert | 1 | | Data Scientist - Mid | 1 | | Data Scientist - Senior | 1 | | Data Scientist - Sr | 1 | | Data Scientist - Supply Chain Risk Management with Security Clearance | 1 | | Data Scientist - Traditional and Generative AI | 1 | | Data Scientist (U.S. Citizen, Secret/Top Secret), Washington D.C. | 1 | | Data Scientist II | 1 | | Data Scientist III | 1 | | Data Scientist Jobs | 1 | | Data Scientist Lead ID | 1 | | Data Scientist SETA (TS/SCI #26-099) | 1 | | Data Scientist, Level 4 with Security Clearance | 1 | | Data Scientist, Mid level | 1 | | Data Scientist, Product Analytics | 1 | | Experienced Data Scientist | 1 | | Junior DoD Data Scientist - Onsite in Alexandria | 1 | Methodology: entities extracted per ad ({hard_skills, tools, methods, responsibilities, seniority}), normalized, counted as DISTINCT ads per entity; report threshold ≥ 3 ads. Headline sections in skills.md/tools.md use the stricter ≥ 20 % threshold.