Q.1
Question 1. Write the statement to install the python connector to connect
MySQL i.e. pymysql
Answer:
To install the pymysql connector, run(type) the following command in
your terminal or command prompt:
pip install pymysql
Q.2
Explain the difference between pivot() and pivot_ table() function?
Answer:
The table below contrasts both reshaping methods:
Feature pivot() pivot_table()
Aggregation Does not support data
aggregation.
Fully supports data
aggregation.
Duplicate
Entries
Raises a ValueError if
duplicate entries exist.
Handles duplicate entries by
aggregating them.
Default
Function No default statistical function. Uses mean as the default
aggregation function.
Q.3
What is sqlalchemy?
Answer:
sqlalchemy is a popular Python SQL toolkit and Object Relational
Mapper (ORM). It provides an efficient interface to interact with relational
databases like MySQL. In Pandas, its create_engine() function is used to
establishes a secure database connection..
Q.4
Can you sort a DataFrame with respect to multiple columns?
Answer:
Yes , we can sort a DataFrame with respect to multiple columns. We
can achieve this by passing a list of multiple column names to the by parameter
of the sort_values() method.
for Example:
df.sort_values(by=['Column1', 'Column2'], ascending=[True, False])
Q.5
What are missing values? What are the strategies to handle them?
Answer:
The Missing values represent blank or unrecorded data points in a
dataset. In Pandas, they are denoted as NaN (Not a Number).
There are two primary strategies to handle them are:
1. Dropping Missing Values: Remove rows or columns containing empty
slots using df.dropna().
2. Filling Missing Values: Replace empty entries with calculated estimates
(like mean or median) using df.fillna().
Q.6
Define the following terms: Median, Standard Deviation and
Variance.
Answer:
Median: The middle value of a sorted, ordered dataset. Calculated using
df.median().
Standard Deviation: A metric showing how much data deviates from the
mean. It is the square root of variance, calculated using df.std().
Variance: The average of squared deviations from the arithmetic mean.
Calculated using df.var().
Q.7
What do you understand by the term MODE? Name the function
which is used to calculate it.
Answer:
The Mode represents the value that appears most frequently in a dataset. A
dataset can have more than one mode. The built-in df.mode() function is used to
calculate mean.
Q.8
Write the purpose of Data aggregation.
Answer:
The purpose of Data aggregation is to combine multiple data records into a
single summary metric. It uses statistical methods (like mean(), sum(), or
count()) to draw quick, high-level insights from large databases.
Q.9
Explain the concept of GROUP BY with help of an example.
Answer:
The groupby() method splits a DataFrame into distinct groups based on column
categories. It follows a Split-Apply-Combine mechanism.
Example: To find the total sales per region:
# Groups by 'Region' and calculates the sum of 'Sales'
df.groupby('Region')['Sales'].sum()
Q.10
Write the steps required to read data from a MySQL database to a
DataFrame.
Answer:
1. Install tools: Run pip install pymysql sqlalchemy.
2. Import modules: Open Python and import pandas and create_engine.
3. Create engine: Establish a link using engine =
create_engine('mysql+pymysql://user:password@host/db').
4. Fetch data: Execute pd.read_sql_query("SELECT * FROM table",
engine) to load records.
Q.11
Explain the importance of reshaping of data with an example.
Answer:
Reshaping data means rearranges rows and columns to make datasets more
organized, readable, and ready for analysis or plotting.
for Example: A teacher tracks marks across columns for multiple terms.
Reshaping converts this wide layout into a clean vertical list, making it simple
to chart individual student progress.
Q.12
Why estimation is an important concept in data analysis?
Answer:
Estimation is vital because real-world data collection often contains missing or
corrupted blocks. Instead of discarding incomplete records—which wastes
valuable information—analysts estimate missing entries using statistical
indicators (like mean or median) to preserve dataset balance.
Q.1
MySQL से जुड़ने के लिए आवश्यक Python connector यानी pymysql को स्थापित (install)
करने का स्टेटमेंट लिखिए।
Answer:
pymysql कनेक्टर को स्थापित करने के लिए अपने टर्मिनल या कमांड प्रॉम्प्ट में निम्नलिखित कमांड
चलाएँ:
pip install pymysql
Q.2
pivot() और pivot_table() फ़ंक्शन के बीच अंतर स्पष्ट कीजिए?
Answer:
डेटा के आकार को बदलने (reshaping) के इन दोनों तरीकों में निम्नलिखित मुख्य अंतर हैं:
विशेषता pivot() pivot_table()
एग्रीगेशन
(Aggregation)
यह डेटा एग्रीगेशन का समर्थन नहीं
करता है। यह डेटा एग्रीगेशन का पूर्ण समर्थन करता है।
डुप्लिकेट एंट्रीज डुप्लिकेट मान होने पर यह
ValueError एरर देता है।
यह डुप्लिकेट मानों को एकत्रित
(aggregate) कर लेता है।
डिफ़ॉल्ट फ़ंक्शन इसमें कोई डिफ़ॉल्ट सांख्यिकीय
फ़ंक्शन नहीं होता।
यह डिफ़ॉल्ट रूप से mean (औसत) फ़ंक्शन
Q.3
sqlalchemy क्या है?
Answer:
sqlalchemy एक लोकप्रिय Python SQL टूलकिट और ऑब्जेक्ट रिलेशनल मैपर (ORM) है। यह
MySQL जैसे रिलेशनल डेटाबेस के साथ इंटरैक्ट करने के लिए एक कुशल इंटरफ़ेस प्रदान करता है। पंडास
(Pandas) में, इसके create_engine() फ़ंक्शन का उपयोग डेटाबेस कनेक्शन स्थापित करने के लिए
किया जाता है।
Q.4
क्या आप एक DataFrame को एक से अधिक कॉलम के आधार पर सॉर्ट कर सकते हैं?
Answer:
हाँ, आप एक DataFrame को एक से अधिक कॉलम के आधार पर सॉर्ट कर सकते हैं। ऐसा करने के लिए,
आपको sort_values() मेथड के by पैरामीटर में कॉलम के नामों की एक लिस्ट पास करनी होती है।
उदाहरण:
df.sort_values(by=['Column1', 'Column2'], ascending=[True, False])
Q.5
मिसिंग वैल्यूज (Missing values) क्या हैं? इनसे निपटने की रणनीतियाँ क्या हैं?
Answer:
जब किसी डेटासेट में किसी वेरिएबल के लिए कोई मान दर्ज नहीं होता, तो उसे मिसिंग वैल्यू (Missing
Value) कहा जाता है। पंडास में इसे NaN (Not a Number) द्वारा दर्शाया जाता है।
इनसे निपटने की दो मुख्य रणनीतियाँ हैं:
1. मिसिंग वैल्यूज को हटाना (Dropping): खाली मानों वाली पंक्तियों या कॉलम को df.dropna()
फ़ंक्शन का उपयोग करके हटा दिया जाता है।
2. मिसिंग वैल्यूज को भरना (Filling): खाली स्थानों को किसी अनुमानित मान (जैसे mean या
median) से बदलने के लिए df.fillna() का उपयोग किया जाता है।
Q.6
निम्नलिखित शब्दों को परिभाषित करें:
माध्यिका (Median), मानक विचलन (Standard Deviation) और विचरण (Variance)।
Answer:
माध्यिका (Median): यह एक व्यवस्थित/क्रमबद्ध डेटासेट का मध्य मान (middle value) होता है। इसे
df.median() द्वारा निकाला जाता है।
मानक विचलन (Standard Deviation): यह यह मापता है कि डेटा अपने माध्य (mean) से
कितना विचलित होता है। यह विचरण का वर्गमूल (square root) है और इसे df.std() द्वारा
निकाला जाता है।
विचरण (Variance): यह अंकगणितीय माध्य से मानों के वर्ग विचलनों का औसत होता है। इसे
df.var() फ़ंक्शन की मदद से मापा जाता है।
Q.7
MODE शब्द से आप क्या समझते हैं? इसे कैलकुलेट करने वाले फ़ंक्शन का नाम लिखिए।
Answer:
मोड (Mode या बहुलक) वह मान है जो किसी डेटासेट में सबसे अधिक बार दिखाई देता है। किसी डेटासेट
में एक से अधिक मोड हो सकते हैं। इसे कैलकुलेट करने के लिए पंडास के इन-बिल्ट df.mode() फ़ंक्शन का
उपयोग किया जाता है।
Q.8
डेटा एग्रीगेशन (Data aggregation) का उद्देश्य लिखिए।
Answer:
डेटा एग्रीगेशन का मुख्य उद्देश्य बहुत सारे डेटा रिकॉर्ड्स को मिलाकर एक एकल संक्षिप्त मान (single
summary metric) में बदलना है। सांख्यिकीय फ़ंक्शंस (जैसे mean(),sum() या count()) का उपयोग
करके बड़े डेटासेट से तुरंत महत्वपूर्ण निष्कर्ष निकाले जा सकते हैं।
Q.9
एक उदाहरण की सहायता से GROUP BY की अवधारणा को समझाइए।
Answer:
groupby() फ़ंक्शन का उपयोग डेटा को किसी विशेष कॉलम के श्रेणियों के आधार पर अलग-अलग समूहों
(groups) में विभाजित करने के लिए किया जाता है। यह Split-Apply-Combine रणनीति पर काम
करता है।
उदाहरण: यदि हमें क्षेत्रवार (Region-wise) कुल बिक्री (Sales) निकालनी हो:
# यह 'Region' के अनुसार डेटा ग्रुप करेगा और 'Sales' का कुल योग निकालेगा
df.groupby('Region')['Sales'].sum()
Q.10
MySQL डेटाबेस से DataFrame में डेटा पढ़ने के लिए आवश्यक चरणों को लिखिए।.
Answer:
1. लाइब्रेरी इंस्टॉल करें: सबसे पहले pip install pymysql sqlalchemy चलाएं।
2. मॉड्यूल इम्पोर्ट करें: पाइथन स्क्रिप्ट में pandas और sqlalchemy से create_engine को
इम्पोर्ट करें।
3. इंजन बनाएं: कनेक्शन सेट करने के लिए लिखें:
4. engine = create_engine('mysql+pymysql://user:password@host/db')
5. डेटा पढ़ें: DataFrame में डेटा लोड करने के लिए pd.read_sql_query("SELECT *
FROM table", engine) का उपयोग करें।.
Q.11
उदाहरण सहित डेटा के रीशेपिंग (Reshaping) के महत्व को समझाइए।
Answer:
डेटा को रीशेप (Reshape) करने का अर्थ है DataFrame की पंक्तियों और कॉलम की बनावट को
बदलना, जिससे डेटा अधिक व्यवस्थित, पढ़ने योग्य और विश्लेषण के अनुकूल हो जाता है।
उदाहरण: एक शिक्षक के पास कई परीक्षाओं के अंकों का चौड़ा (wide) टेबल है। उसे रीशेप करके एक
लंबवत (vertical) सूची में बदलने से प्रत्येक छात्र के प्रदर्शन का ग्राफ बनाना या विश्लेषण करना आसान हो
जाता है।
Q.12
डेटा विश्लेषण (Data analysis) में अनुमान (Estimation) एक महत्वपूर्ण अवधारणा क्यों है?
Answer:
वास्तविक दुनिया के डेटासेट में अक्सर अधूरा या मिसिंग डेटा होता है। यदि हम उन सभी अधूरी पंक्तियों
को हटा देंगे, तो बहुत सी मूल्यवान जानकारी नष्ट हो जाएगी। इसलिए, डेटा के संतुलन को बनाए रखने
और सही निष्कर्ष निकालने के लिए अनुमान (Estimation) की मदद से खाली जगहों पर माध्य (mean)
या माध्यिका (median) जैसे मान भर दिए जाते हैं।