Обучение аналитиков данных Cloudera
Cloudera Data Analyst Training
Cloudera University’s four-day Data Analyst Training course will teach you to apply traditional
data analytics and business intelligence skills to big data. This course presents the tools data
professionals need to access, manipulate, transform, and analyze complex data sets using SQL
and familiar scripting languages.
Advance Your Ecosystem Expertise
Apache Hive makes transformation and analysis of complex, multi-structured data scalable in Cloudera environments. Apache Impala enables real-time interactive analysis of the data stored in Hadoop using a native SQL environment. Together, they make multi-structured data accessible to analysts, database administrators, and others without Java programming expertise.
What to Expect
Through instructor-led discussion and interactive, hands-on exercises, participants will navigate the ecosystem, learning:
- How the open source ecosystem of big data tools addresses challenges not met by traditional RDBMSs
- Using Apache Hive and Apache Impala to provide SQL access to data
- Hive and Impala syntax and data formats, including functions and subqueries
- Create, modify, and delete tables, views, and databases; load data; and store results of queries
- Create and use partitions and dierent file formats
- Combining two or more datasets using JOIN or UNION, as appropriate
- What analytic and windowing functions are, and how to use them Store and query complex or nested data structures
- Process and analyze semi-structured and unstructured data
- Techniques for optimizing Hive and Impala queries
- Extending the capabilities of Hive and Impala using parameters, custom file formats and SerDes, and external scripts
- How to determine whether Hive, Impala, an RDBMS, or a mix of these is best for a given task
Audience & Prerequisites
This course is designed for data analysts, business intelligence specialists, developers, system architects, and database administrators. Some knowledge of SQL is assumed, as is basic Linux command-line familiarity. Prior knowledge of Apache Hadoop is not required.
Upon completion of the course, attendees are encouraged to continue their study and register for the CCA Data Analyst exam. Certification is a great dierentiator. It helps establish you as a leader in the field, providing employers and customers with tangible evidence of your skills and expertise