Files
skillfactor-pipeline/adapters/claude/dist/data-scientist/competences/manage-data.md
skillfactor-pipeline 5c03e07ded fix(homepage): chat example in English, simpler case, real chat bubbles
Simpler story (slide decks unread -> three-bullet status email), proper
Claude-chat look with avatars and bubbles on both sides. Canonical trigger
prompt switched to English everywhere (homepage, architecture, both
adapter generators rebuilt) - V7 consistency green, all checks pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PDKeXvpT6tENSvyQGLV1Uq
2026-07-10 06:01:43 +02:00

1.2 KiB
Raw Blame History

esco_uri, esco_label, relation, onet_soc, source, confidence, qa_count, generator, generated
esco_uri esco_label relation onet_soc source confidence qa_count generator generated
http://data.europa.eu/esco/skill/9ff9db9d-d14b-426e-83f3-e7449af6c79f manage data optional 15-2051.00 model-knowledge high 0 gemma3:27b (prompt-designed and spot-checked by Claude) 2026-07-10

manage data — Data Scientist

For a Data Scientist, 'managing data' isnt just about storage; it's the majority of project time. It means taking raw, messy inputs think website logs, sensor readings, customer databases and transforming them into analysis-ready datasets. Daily tasks include profiling (understanding distributions & anomalies), cleaning (handling missing values, correcting errors), and feature engineering (creating new variables). You'll be writing code—primarily Python with libraries like Pandas, NumPy, and potentially Spark for large datasets—to parse different formats (CSV, JSON, SQL databases) and standardize data types. Its less about administration in a sysadmin sense, and more about data wrangling to ensure analytical validity.

Weitere Anreicherung

Stage-2 source for future practitioner grounding: arXiv cs.LG/stat.ML (ML preprints).