en
Feedback
Data Science & Machine Learning

Data Science & Machine Learning

Open in Telegram

The first channel on Telegram that offers exciting questions, answers, and tests in data science, artificial intelligence, machine learning, and programming languages. For promotions: @love_data

Show more

📈 Analytical overview of Telegram channel Data Science & Machine Learning

Channel Data Science & Machine Learning (@datascienceinterviews) in the English language segment is an active participant. Currently, the community unites 27 639 subscribers, ranking 6 938 in the Education category and 14 632 in the India region.

📊 Audience metrics and dynamics

Since its creation on невідомо, the project has demonstrated rapid growth, gathering an audience of 27 639 subscribers.

According to the latest data from 31 August, 2026, the channel demonstrates stable activity. Although there has been a change in the number of participants by 194 over the last 30 days and by 15 over the last 24 hours, overall reach remains high.

  • Verification status: Not verified
  • Engagement rate (ER): The average audience engagement rate is 2.32%. Within the first 24 hours after publication, content typically collects 0.48% reactions from the total number of subscribers.
  • Post reach: On average, each post receives 641 views. Within the first day, a publication typically gains 133 views.
  • Reactions and interaction: The audience actively supports content: the average number of reactions per post is 5.
  • Thematic interests: Content is focused on key topics such as insidead, mining, pinix, learning, neo.

📝 Description and content policy

The author describes the resource as a platform for expressing subjective opinions:
The first channel on Telegram that offers exciting questions, answers, and tests in data science, artificial intelligence, machine learning, and programming languages. For promotions: @love_data

Thanks to the high frequency of updates (latest data received on 01 September, 2026), the channel maintains relevance and a high level of publication reach. Analytics show that the audience actively interacts with content, making it an important point of influence in the Education category.

27 639
Subscribers
+1524 hours
+547 days
+19430 days
Posts Archive
Who's here?  We've asked for a free link to a paid channel, for our subs. x2-x3 Signals here 👉 CLICK HERE TO JOIN 👈 👉 CLICK HERE TO JOIN 👈 👉 CLICK HERE TO JOIN 👈 ❗️JOIN FAST! FIRST 1000 SUBS WILL BE ACCEPTED

Ad 👇👇

THE MOST PRIVATE GROUP №1 ❌ They are robbing Crypto Exchanges for Millions of dollars! Yesterday profit = 50,000$+ 👉 https://t.me/+VubRJjjSR_o4MzI1 👉 https://t.me/+VubRJjjSR_o4MzI1 👉 https://t.me/+VubRJjjSR_o4MzI1 JOIN FAST! First 1000 subs will be accepted! 👀🚀

#ad

Data Science Interview Questions and Answers.pdf1.76 MB

1. What are Different Kernels in SVM? Linear kernel - used when data is linearly separable. Polynomial kernel - When you have discrete data that has no natural notion of smoothness. Radial basis kernel - Create a decision boundary able to do a much better job of separating two classes than the linear kernel. Sigmoid kernel - used as an activation function for neural networks. 2. What is Cross-Validation? Cross-validation is a method of splitting all your data into three parts: training, testing, and validation data. Data is split into k subsets, and the model has trained on k-1of those datasets. The last subset is held for testing. This is done for each of the subsets. This is k-fold cross-validation. Finally, the scores from all the k-folds are averaged to produce the final score. 3. List the different types of relationships in SQL. One-to-One - This can be defined as the relationship between two tables where each record in one table is associated with the maximum of one record in the other table. One-to-Many & Many-to-One - This is the most commonly used relationship where a record in a table is associated with multiple records in the other table. Many-to-Many - This is used in cases when multiple instances on both sides are needed for defining a relationship. Self-Referencing Relationships - This is used when a table needs to define a relationship with itself. 4. What Are the Data Types Supported in Tableau? Following data types are supported in Tableau: Text (string) values Date values Date and time values Numerical values Boolean values (relational only) Geographical values (used with maps) ENJOY LEARNING 👍👍

THE MOST PRIVATE GROUP №1 ❌ They are robbing Crypto Exchanges for Millions of dollars! Yesterday profit = 50,000$+ 👉 https://t.me/+VubRJjjSR_o4MzI1 👉 https://t.me/+VubRJjjSR_o4MzI1 👉 https://t.me/+VubRJjjSR_o4MzI1 JOIN FAST! First 1000 subs will be accepted! 👀🚀

#ad

1. What do you understand by the term silhouette coefficient? The silhouette coefficient is a measure of how well clustered together a data point is with respect to the other points in its cluster. It is a measure of how similar a point is to the points in its own cluster, and how dissimilar it is to the points in other clusters. The silhouette coefficient ranges from -1 to 1, with 1 being the best possible score and -1 being the worst possible score. 2. What is the difference between trend and seasonality in time series? Trends and seasonality are two characteristics of time series metrics that break many models. Trends are continuous increases or decreases in a metric’s value. Seasonality, on the other hand, reflects periodic (cyclical) patterns that occur in a system, usually rising above a baseline and then decreasing again. 3. What is Bag of Words in NLP? Bag of Words is a commonly used model that depends on word frequencies or occurrences to train a classifier. This model creates an occurrence matrix for documents or sentences irrespective of its grammatical structure or word order. 4. What is the difference between bagging and boosting? Bagging is a homogeneous weak learners’ model that learns from each other independently in parallel and combines them for determining the model average. Boosting is also a homogeneous weak learners’ model but works differently from Bagging. In this model, learners learn sequentially and adaptively to improve model predictions of a learning algorithm

1. What are the uses of using RNN in NLP? The RNN is a stateful neural network, which means that it not only retains information from the previous layer but also from the previous pass. Thus, this neuron is said to have connections between passes, and through time. For the RNN the order of the input matters due to being stateful. The same words with different orders will yield different outputs. RNN can be used for unsegmented, connected applications such as handwriting recognition or speech recognition. 2. How to remove values to a python array? Ans: Array elements can be removed using pop() or remove() method. The difference between these two functions is that the former returns the deleted value whereas the latter does not. 3. What are the advantages and disadvantages of views in the database? Answer: Advantages of Views: As there is no physical location where the data in the view is stored, it generates output without wasting resources. Data access is restricted as it does not allow commands like insertion, updation, and deletion. Disadvantages of Views: The view becomes irrelevant if we drop a table related to that view. Much memory space is occupied when the view is created for large tables. 4. How to create a calculated field in Tableau? Click the drop down to the right of Dimensions on the Data pane and select “Create > Calculated Field” to open the calculation editor. Name the new field and create a formula. ENJOY LEARNING 👍👍

THE MOST PRIVATE GROUP №1 ❌ They are robbing Crypto Exchanges for Millions of dollars! Yesterday profit = 50,000$+ 👉 https://t.me/+5YRwSjrwkCIwNDE1 👉 https://t.me/+5YRwSjrwkCIwNDE1 👉 https://t.me/+5YRwSjrwkCIwNDE1 Go fast! Only the first 1000 subs will be accepted! 👀🚀

Coding and Aptitude Round before interview Coding challenges are meant to test your coding skills (especially if you are applying for ML engineer role). The coding challenges can contain algorithm and data structures problems of varying difficulty. These challenges will be timed based on how complicated the questions are. These are intended to test your basic algorithmic thinking. Sometimes, a complicated data science question like making predictions based on twitter data are also given. These challenges are hosted on HackerRank, HackerEarth, CoderByte etc. In addition, you may even be asked multiple-choice questions on the fundamentals of data science and statistics. This round is meant to be a filtering round where candidates whose fundamentals are little shaky are eliminated. These rounds are typically conducted without any manual intervention, so it is important to be well prepared for this round. Sometimes a separate Aptitude test is conducted or along with the technical round an aptitude test is also conducted to assess your aptitude skills. A Data Scientist is expected to have a good aptitude as this field is continuously evolving and a Data Scientist encounters new challenges every day. If you have appeared for GMAT / GRE or CAT, this should be easy for you. Resources for Prep: For algorithms and data structures prep,Leetcode and Hackerrank are good resources. For aptitude prep, you can refer to IndiaBixand Practice Aptitude. With respect to data science challenges, practice well on GLabs and Kaggle. Brilliant is an excellent resource for tricky math and statistics questions. For practising SQL, SQL Zoo and Mode Analytics are good resources that allow you to solve the exercises in the browser itself. Things to Note: Ensure that you are calm and relaxed before you attempt to answer the challenge. Read through all the questions before you start attempting the same. Let your mind go into problem-solving mode before your fingers do! In case, you are finished with the test before time, recheck your answers and then submit. Sometimes these rounds don’t go your way, you might have had a brain fade, it was not your day etc. Don’t worry! Shake if off for there is always a next time and this is not the end of the world.

❓ Question 2: What is a z-score? 1. A standardized value that indicates the number of standard deviations an observation is from the mean. 2. The range between the highest and lowest values in a set of data. 3. A measure of the spread of a set of data. 4. A measure of central tendency of a set of data. ✅ Correct Response: 1 Explanation: In statistics, a z-score is a standardized value that indicates the number of standard deviations an observation is from the mean of a set of datIt is used to compare values from different normal distributions and to calculate probabilities. https://t.me/DataScienceInterviews

❓ Question 1: What is the difference between a population and a sample in statistics? 1. A population is a subset of a sample. 2. A sample is a subset of a population. 3. A population is a larger group, while a sample is a smaller group. 4. A sample is a group that is more representative than a population. ✅ Correct Response: 2 Explanation: In statistics, a population is the entire group of individuals, objects, or events that we are interested in studying, while a sample is a smaller subset of the population that is selected for study. Samples are often used when it is not feasible or practical to study the entire population https://t.me/DataScienceInterviews

1.How will you handle missing values in data? There are several ways to handle missing values in the given data- 1.Dropping the values 2.Deleting the observation (not always recommended). 3.Replacing value with the mean, median and mode of the observation. 4.Predicting value with regression 5.Finding appropriate value with clustering 2. What is SVM? Can you name some kernels used in SVM? SVM stands for support vector machine. They are used for classification and prediction tasks. SVM consists of a separating plane that discriminates between the two classes of variables. This separating plane is known as hyperplane. Some of the kernels used in SVM are – Polynomial Kernel Gaussian Kernel Laplace RBF Kernel Sigmoid Kernel Hyperbolic Kernel 3.What is market basket analysis? Market Basket Analysis is a modeling technique based upon the theory that if you buy a certain group of items, you are more (or less) likely to buy another group of items. 4.What is the benefit of batch normalization? The model is less sensitive to hyperparameter tuning. High learning rates become acceptable, which results in faster training of the model. Weight initialization becomes an easy task. Using different non-linear activation functions becomes feasible. Deep neural networks are simplified because of batch normalization. It introduces mild regularisation in the network.

1. Explain character-manipulation functions? Explains its different types in SQL. Change, extract, and edit the character string using character manipulation routines. The function will do its action on the input strings and return the result when one or more characters and words are supplied into it. The character manipulation functions in SQL are as follows: A) CONCAT (joining two or more values): This function is used to join two or more values together. The second string is always appended to the end of the first string. B) SUBSTR: This function returns a segment of a string from a given start point to a given endpoint. C) LENGTH: This function returns the length of the string in numerical form, including blank spaces. D) INSTR: This function calculates the precise numeric location of a character or word in a string. E) LPAD: For right-justified values, it returns the padding of the left-side character value. F) RPAD: For a left-justified value, it returns the padding of the right-side character value. G) TRIM: This function removes all defined characters from the beginning, end, or both ends of a string. It also reduced the amount of wasted space. H) REPLACE: This function replaces all instances of a word or a section of a string (substring) with the other string value specified. 2. How Do You Calculate the Daily Profit Measures Using LOD? LOD expressions allow us to easily create bins on aggregated data such as profit per day. Scenario: We want to measure our success by the total profit per business day. Create a calculated field named LOD - Profit per day and enter the formula: FIXED [Order Date] : SUM ([Profit]) Create another calculated field named LOD - Daily Profit KPI and enter the formula: IF [LOD - Profit per day] > 2000 then “Highly Profitable.” ELSEIF [LOD - Profit per day] <= 0 then “Unprofitable” ELSE “Profitable” END To calculate daily profit measure using LOD, follow these steps to draw the visualization: Bring YEAR(Order Date) and MONTH(Order Date) to the Columns shelf Drag Order Id field to Rows shelf. Right-click on it, select Measure and click on Count(Distinct) Drag LOD - Daily Profit KPI to the Rows shelf Bring LOD - Daily Profit KPI to marks card and change mark type from automatic to area. 3. What are Superkey and candidate key? A super key may be a single or a combination of keys that help to identify a record in a table. Know that Super keys can have one or more attributes, even though all the attributes are not necessary to identify the records. A candidate key is the subset of Superkey, which can have one or more than one attributes to identify records in a table. Unlike Superkey, all the attributes of the candidate key must be helpful to identify the records. Note that all the candidate keys can be Super keys, but all the super keys cannot be candidate keys. 4.What is Database Cardinality? Database Cardinality denotes the uniqueness of values in the tables. It supports optimizing query plans and hence improves query performance. There are three types of database cardinalities in SQL, as given below: Higher Cardinality Normal Cardinality Lower Cardinality

Do you enjoy reading this channel? Perhaps you have thought about placing ads on it? To do this, follow three simple steps: 1) Sign up: https://telega.io/c/DataScienceInterviews 2) Top up the balance in a convenient way 3) Create an advertising post If the topic of your post fits our channel, we will publish it with pleasure.

👉✔️Here are Data Analytics-related questions along with their answers: 1.Question: What is the purpose of exploratory data analysis (EDA)? Answer: EDA is used to analyze and summarize data sets, often through visual methods, to understand patterns, relationships, and potential outliers. 2. Question: What is the difference between supervised and unsupervised learning? Answer: Supervised learning involves training a model on a labeled dataset, while unsupervised learning deals with unlabeled data to discover patterns without explicit guidance. 3.Question: Explain the concept of normalization in the context of data preprocessing. Answer: Normalization scales numeric features to a standard range, preventing certain features from dominating due to their larger scales. 4. Question: What is the purpose of a correlation coefficient in statistics? Answer: A correlation coefficient measures the strength and direction of a linear relationship between two variables, ranging from -1 to 1. 5. Question: What is the role of a decision tree in machine learning? Answer: A decision tree is a predictive model that maps features to outcomes by recursively splitting data based on feature conditions. 6. Question: Define precision and recall in the context of classification models. Answer: Precision is the ratio of correctly predicted positive observations to the total predicted positives, while recall is the ratio of correctly predicted positive observations to all actual positives. 7. Question: What is the purpose of cross-validation in machine learning? Answer: Cross-validation assesses a model's performance by dividing the dataset into multiple subsets, training the model on some, and testing it on others, helping to evaluate its generalization ability. 8. Question: Explain the concept of a data warehouse. Answer: A data warehouse is a centralized repository that stores, integrates, and manages large volumes of data from different sources, providing a unified view for analysis and reporting. 9. Question: What is the difference between structured and unstructured data? Answer: Structured data is organized and easily searchable (e.g., databases), while unstructured data lacks a predefined structure (e.g., text documents, images). 10. Question: What is clustering in machine learning? Answer: Clustering is a technique that groups similar data points together based on certain features, helping to identify patterns or relationships within the data.

1. What are the common problems that data analysts encounter during analysis? The common problems steps involved in any analytics project are: Handling duplicate data Collecting the meaningful right data at the right time Handling data purging and storage problems Making data secure and dealing with compliance issues 2. Explain the Type I and Type II errors in Statistics? In Hypothesis testing, a Type I error occurs when the null hypothesis is rejected even if it is true. It is also known as a false positive. A Type II error occurs when the null hypothesis is not rejected, even if it is false. It is also known as a false negative. 3. What’s the F1 score? How would you use it? The F1 score is a measure of a model’s performance. It is a weighted average of the precision and recall of a model, with results tending to 1 being the best, and those tending to 0 being the worst. 4. Name an example where ensemble techniques might be useful? Ensemble techniques use a combination of learning algorithms to optimize better predictive performance. They typically reduce overfitting in models and make the model more robust (unlikely to be influenced by small changes in the training data). You could list some examples of ensemble methods (bagging, boosting, the “bucket of models” method) and demonstrate how they could increase predictive power. ————————————————————-