Introduction to Data Science (DM1)

Databases, NoSQL and Big Data

This course introduces the field of data science, showing what it covers from exploring data and creating visualizations to applying basic machine learning methods, and interpreting and evaluating results for practical use.

No programming skills are required; the hands-on part runs on a cloud platform (BigML) or on popular desktop tools like RapidMiner and Weka. The course covers data preparation, basic model evaluation, and a practical NLP example.

THIS TRAINING COURSE WILL HELP YOU:

  • Understand core data science concepts and terminology
  • Perform exploratory data analysis and create clear visualizations
  • Prepare and clean data for modeling, including sampling
  • Apply basic machine learning techniques and evaluate models

WHO SHOULD ATTEND?

  • Business analysts and domain experts seeking data skills
  • Professionals without a programming background
  • Decision-makers who need to interpret data outputs
  • Students preparing for roles in analytics or data science

COURSE LOCATION AND AVAILABLE DATES



Public courses are usually delivered in Czech, but this course is also available in English. We can arrange private training for your team online, at your premises or in our classrooms, and tailor the content to your needs.

For groups of around 4 or more participants, private training can already be comparable in price to booking individual places on a public course. Send us your requirements and we’ll recommend the best format and provide an exact quote.

Request training in English

Course content:

Hide details
  • Introduction and key concepts (data science vs. machine learning vs. artificial intelligence vs. data mining)
  • Typical workflow for analytical projects and stages of data analytics
  • Data, data types, and data quality
  • Exploratory data analysis and data visualization
  • Tools for data science and common choices
    1. Local desktop tools (on a local computer)
    2. R and Python languages (for Python, introduction to core libraries)
    3. Cloud platforms
  • Data preparation
    1. Selection
    2. Data cleaning
    3. Data transformation (value grouping, discretization, derived columns, …)
    4. Sampling
  • Machine learning techniques
    1. Linear regression
    2. Classification tasks – logistic regression, decision trees, neural networks, Bayesian approaches
    3. Clustering
    4. Association rules
    5. Anomaly detection
  • Interpreting results and model evaluation
  • Natural language processing with a practical example
  • State of the art in data science, machine learning, and artificial intelligence
Prerequisites:
Basic computer user skills and basic statistics.
Schedule:
2 days (9:00-17:00)

Training and learning environment