Big Data and Analytics PYQ 2023-24 AKTU (KCS061) Question Paper
AKTU · BTECH · Semester 6 · Big Data and Analytics · Session 2023-24 · PYQ
AKTU Big Data and Analytics (KCS061) previous year question paper 2023-24 for B.Tech Semester 6. Covers Introduction to Big Data, Hadoop and MapReduce, HDFS…
Open the interactive reader to study this resource on AcademicArk.
Big Data and Analytics AKTU syllabus
- Unit 1: Introduction to Big Data
- Unit 2: Hadoop and MapReduce
- Unit 3: HDFS and Hadoop Environment
- Unit 4: Hadoop Ecosystem YARN NoSQL Spark Scala
- Unit 5: Hadoop Ecosystem Frameworks Pig Hive HBase
Questions in Big Data and Analytics AKTU PYQ 2023-24
- Q1a. What are the different types of digital data commonly encountered in Big Data applications? Provide examples of structured, semi-structured, and unstructured data. (2 marks, 2023-24)
- Q1b. What constitutes a Big Data platform? (2 marks, 2023-24)
- Q1c. What is Hadoop Streaming? (2 marks, 2023-24)
- Q1d. Discuss the data formats commonly used in Hadoop environments. (2 marks, 2023-24)
- Q1e. Describe the concepts of file sizes, block sizes, and block abstraction in HDFS. (2 marks, 2023-24)
- Q1f. What are the benefits and challenges of using HDFS for distributed storage and processing? (2 marks, 2023-24)
- Q1g. What are the characteristics and use cases for schedulers such as Fair Scheduler and Capacity Scheduler? (2 marks, 2023-24)
- Q1h. What is YARN? (2 marks, 2023-24)
- Q1i. What is Apache Pig? (2 marks, 2023-24)
- Q1j. Describe the Grunt shell in Apache Pig. (2 marks, 2023-24)
- Q2a. Distinguish between data analysis and reporting in the context of Big Data. How does advanced analytics go beyond traditional reporting to uncover hidden patterns, trends, and correlations in data? (10 marks, 2023-24)
- Q2b. Explain Apache Hadoop and its role in big data processing. What are the core components of the Apache Hadoop ecosystem, and how do they work together to enable distributed data storage and processing? (10 marks, 2023-24)
- Q2c. Explain the core concepts of HDFS, including NameNode, DataNode, and the file system namespace. How do these components work together to manage data storage and replication in Hadoop clusters? (10 marks, 2023-24)
- Q2d. Define NoSQL databases. What are the key characteristics and benefits of NoSQL databases compared to traditional relational databases? (10 marks, 2023-24)
- Q2e. Provide an overview of Apache Hive architecture and its components. How does Hive translate SQL-like queries into MapReduce jobs for data processing in Hadoop? (10 marks, 2023-24)
- Q3a. Describe the "5 Vs" of Big Data. What do each of these terms represent in the context of Big Data, and why are they essential considerations for data management and analysis? (10 marks, 2023-24)
- Q3b. Provide examples of real-world applications where Big Data analytics have been instrumental. How do industries such as healthcare, finance, e-commerce, and transportation leverage Big Data to gain insights and create value? (10 marks, 2023-24)
- Q4a. Describe the Hadoop Distributed File System (HDFS). How does HDFS manage the storage and replication of data across a distributed cluster of machines? (10 marks, 2023-24)
- Q4b. Discuss the process of developing a MapReduce application. What are the key steps involved in writing, testing, and deploying a MapReduce program? (10 marks, 2023-24)
- Q5a. Explain how HDFS stores, reads, and writes files. Describe the sequence of operations involved in storing a file in HDFS, retrieving data from HDFS, and writing data to HDFS. (10 marks, 2023-24)
- Q5b. Describe the considerations for deploying Hadoop in a cloud environment. What are the advantages and challenges of running Hadoop clusters on cloud platforms like Amazon Web Services (AWS), Microsoft Azure, and Google Cloud Platform (GCP)? (10 marks, 2023-24)
- Q6a. Explain the operations for creating, updating, and deleting documents in MongoDB. What are the MongoDB CRUD operations, and how are they used to manipulate data in collections? (10 marks, 2023-24)
- Q6b. Discuss Resilient Distributed Datasets (RDDs) in Spark. What are RDDs, and how do they enable fault-tolerant and distributed data processing in Spark applications? (10 marks, 2023-24)
- Q7a. Introduce the concepts of HBase and its role in the Hadoop ecosystem. How does HBase differ from traditional relational databases, and what advantages does it offer for storing and accessing large-scale data? (10 marks, 2023-24)
- Q7b. Discuss the HiveQL language used in Apache Hive. How does HiveQL support SQL-like syntax for defining tables, querying data, and performing data manipulation operations? (10 marks, 2023-24)
AKTU paper codes: BCDS601, BCS061, KDS601, KCS061, NIT067
Big Data and Analytics previous year papers
More Big Data and Analytics resources
Browse all notes · Semester 6 notes