Showing posts with label regression. Show all posts
Showing posts with label regression. Show all posts

Monday, 21 July 2025

Basic Research Methods and Statistical Data Analysis | Book Publisher International

 

Research is an integral component of scientific enquiry and involves the objective investigation of phenomena. Statistical analyses provide an indispensable tool for conducting unbiased testing of scientific hypotheses. While qualitative research uses narratives, phenomenology, ethnographies, grounded theory and case studies in social or behavioural studies, quantitative approaches involve designed experiments and statistical analyses of instrument-based, performance, observational or attitude data. Valid statistical analyses rely on probability sampling to ensure random and unbiased collection of the data.  These include simple random sampling, systematic sampling, stratified random sampling and cluster sampling. Key concepts in statistics are central tendency, which is reflected by the mean, median and mode of a data set. Range, variance and standard deviation indicate the spread and variability of the data. The definition and classification of variables in a study is important as it specifies the type of data being collected, the statistical models that are appropriate, and the statistical test to be used.

 

Probability theory involves making predictions about the chances of the occurrence of events based on assumptions about the underlying probability process. Probability mass functions describe the possible outcomes and their probabilities for discrete random variables, while probability density functions are used to summarise the information in probability distributions for continuous random variables. Binomial and Poisson distributions are examples of discrete probability distributions, whereas t-distribution, normal distribution, Chi-square distribution and F-distribution are continuous distributions.  Degrees of freedom in statistics indicate the possible number for which a factor or parameter is “free to vary” and is usually one less than the number of variables in each source factor. Exploratory data analysis, done before the actual statistical analyses, helps researchers to understand the nature of the data and to choose the best methods to analyse it. Four types of EDA are univariate non-graphical, univariate graphical, multivariate non-graphical and multivariate graphical techniques.

 

Non-parametric tests are methods of analysing data that do not require the data to follow a distribution. They are generally used when the data do not meet the required assumptions for applying the parametric test, such as the t-test or one-way analysis of variance. The Chi-square goodness of fit test is used to evaluate the probability of an expected outcome when it can be approximated by a Chi-square distribution, and is commonly used for categorical data. The chi-square test of independence is used to determine whether two categorical variables are dependent upon each other or not. The Wilcoxon signed-rank test is used to compare two populations when the assumptions for the t-test do not hold. The Wilcoxon signed rank test can be used as a substitute for the paired t-test and employs both the magnitudes and signs of the differences between pairs of measurements that are ranked and compared to a fixed value D0. The Kruskal-Wallis test is an extension of the sum rank test used to compare more than two populations, and therefore, is an alternative to the one-way analysis of variance.

 

True experiments require random assignment of treatments to subjects, and the tests used assume that the data to be analysed are continuous and follow a normal (Gaussian) distribution. The t-test, completely randomised design, two-way analysis of variance, factorial experiments and split-plot or split-block designs are common experimental designs used in true experiments. When a statistical test establishes a significant difference among the treatments, it may be wished to further determine which treatments differ significantly from the others and which do not. Fisher’s Least Significant Difference and Tukey’s W procedure are two popular methods for conducting multiple means comparisons.

 

Regression analysis is done to establish the relationship between a dependent variable and one or more independent variables. Linear regression is used when a dependent variable is related to a single independent variable. The least-squares method minimises the sum of squares of the errors of prediction for fitting a straight line to the data set. When a straight line does not adequately represent the relationship between a dependent and independent variable, non-linear regression models may be used. These can include exponential, power or polynomial equation fitting of the data set. Multiple regression entails using a polynomial model relating a dependent variable to a set of quantitative independent variables.

 

Author(s) Details

Roshan Man Bajracharya
Kathmandu University, Nepal.

Please see the book here:- https://doi.org/10.9734/bpi/mono/978-81-990309-3-0

Monday, 30 June 2025

Precision and Performance: Regression Methods in Data Processing | Chapter 3 | Engineering Research: Perspectives on Recent Advances Vol. 8


Regression analysis is a fundamental statistical technique used for modelling relationships between multiple variables, playing a significant role in predictive analytics and artificial intelligence. It helps in evaluating dependencies between a dependent variable and one or more independent variables, making it essential for forecasting outcomes. This study focuses on the application of regression models to analyse vehicle dynamics. By utilising data such as traffic velocity, road gradient, and actual speed, we aim to predict a vehicle’s velocity profile. Various regression techniques—including linear regression, multivariate linear regression, and nonlinear regression—are examined to determine their effectiveness in data processing. The core objective of this research is to develop reusable functions for each model, eliminating dependence on predefined programming functions, while enabling effective visualisation of data for optimal model selection. One advantage of a regression model over factor or cluster analysis is that the regression model can be used to obtain an estimate of the actual amount of change in a dependent variable that occurs as a result of a change in an independent variable. This study only focuses on simple linear regression; future work should focus on the use of multiple regression to capture more complex relationships.

 

Author (s) Details

Sandhya Tatekalva
Department of Computer Science, S.V. University, Tirupati, India.

 

Please see the book here:- https://doi.org/10.9734/bpi/erpra/v8/5675

 

Saturday, 21 June 2025

Application of Basic Regression Analysis Models: A Statistical Approach | Chapter 3 | Mathematics and Computer Science: Contemporary Developments Vol. 4

 

Regression analysis is one of the well-liked statistical methods which is used for prediction analysis. In statistical data analysis, one usually needs to begin an association between the various parameters in a data set. This association is crucial for prediction and analysis. So, regression is one of the techniques for prediction analysis and data mining tasks. Each one has its own sense. To construct future predictions, Regression analysis comprises fitting the right model relating to the inclined data set. These techniques vary in terms of the type of response variable, explanatory variable and distribution. This work is mainly focused on the dissimilar types of regression techniques premeditated for various types of analysis and which types of regression are used in the context of different data sets. The study discussed the four types of regression models such as Linear Regression, Polynomial Regression, Partial Least Square Regression and Principal Component Regression, and Support Vector Regression in detail. Although the polynomial regression model agrees on a non-linear association between the response variable and explanatory variable, still it is observed as linear regression since its regression coefficients a0,a1,a2,........an are linear.

 

Author (s) Details

R. Reka
Department of Statistics, Sri Sarada College for Women (Autonomous), Salem, India.

 

Please see the book here:- https://doi.org/10.9734/bpi/mcscd/v4/1317

Wednesday, 19 February 2025

Predicting Rice Crop Yields in Prayagraj: A Comparison of Regression Techniques and Neural Networks | Chapter 3 | Geography, Earth Science and Environment: Research Highlights Vol. 5

This study investigates rice crop yield prediction in Prayagraj District, Uttar Pradesh, using twenty-nine years of historical weather and crop yield data from 1991 to 2019. The analysis divided the dataset into a calibration segment comprising 90% of the data over 26 years, with the remaining 10% used for validation. In our approach, 75.9% of the data was used to train an artificial neural network (ANN) model, and 24.1% was employed for testing and validation to ensure 100% model efficiency evaluation. Techniques such as stepwise linear regression and neural networks were applied to predict rice yields. The performance of these models was measured primarily by the Normalized Root Mean Squared Error (nRMSE), with the regression-based model achieving superior performance, indicated by the lowest nRMSE values. The study also noted the critical role of Bright Sunshine Hours, which demonstrated a significant predictive power with an nRMSE of 0.00025 and a coefficient of determination of 0.94.

 

Author (s) Details

 

Nilesh Kumar Singh
Sam Higginbottom University of Agriculture Technology and Sciences, Allahabad, Uttar Pradesh, India.

 

Shraddha Rawat
Sam Higginbottom University of Agriculture Technology and Sciences, Allahabad, Uttar Pradesh, India.

 

Please see the book here:- https://doi.org/10.9734/bpi/geserh/v5/4249

Saturday, 30 December 2023

Customer Segmentation Using R-Studio | Chapter 7 | Research and Applications Towards Mathematics and Computer Science Vol. 7

Client segmentation is the process of dividing consumers into groups based on common traits. Segmentation is one of the key facets of business decision group providing support to members. In order to grow the business cleverly in competitive market, identification of potential client should be done up-to-date. It can help a business to better understand allure target audience. Separation helps marketers to be more efficient in conditions of time, money and added resources. customer separation allows companies to discover about their customers. Regression reasoning plays a key role in customer segmentation. By accumulating the previous sales, device, customers information individual can improve their company tumor on a wider range and R studio is secondhand for data visualization.

Author(s) Details:

R. Srilatha,
Department of Mathematics, VNRVJIET, Hyderabad, India.

M. Sridevi,
Department of Mathematics, Koneru Lakshmaiah Education Foundation, Hyderabad, India.

Please see the link here: https://stm.bookpi.org/RATMCS-V7/article/view/12878


Wednesday, 27 September 2023

Time-Dependent Cox Regression Analysis of Rock Salt | Chapter 7 | Advances and Challenges in Science and Technology Vol. 2

 The main objective concerning this study is to analyse and anticipate the potential risk of rock salt damage using the Cox equivalent hazard model. Analysed according to the belief of continuous mechanism medium rock massif is a natural environment troublesome to know, so the forecasts refer to its behaviour are approximate and changeable. Such an assertion enhances a certitude, even developing from reality; results of the four fundamental features of mountain system are multiple and include really troublesome issues in mining field and mainly, in the constructions field. The concept and the habit to assess the stability of secret construction includes understanding its phenomenological development and doubtless the personality of rocks behaviour in massif. Most of moment of truth, the analysis of the behaviour of the rocks and the chance of their breaking requires concern of the simultaneous belongings of several determinants. Rocks behave otherwise over time, and in any position, there is a feasibility that the break will manifest itself. Many times, we do not have complete information for the study of the resistance period of the rocks. To carry out aforementioned an analysis, we projected the use of a regressive means, that is, the Cox equivalent hazard model. By applying the Cox model or equivalent hazard regression, we have the feasibility of predicting the relative risk (potential risk) of the incident of rock breakage based on various variables. Risk (hazard) can have a different alternative, either in the sense that it can increase in time or possibly decrease; if over time skilled is no variable or determinant that determines an progress towards breaking the rock, then the risk of breaking is lower. Based on the standard scheme of rocks behaviour habit to failure, we are intend a statistical reasoning method of results acquired by laboratory tests using Cox reversion. This method offer the chance to provide quantitatively and qualitatively moment of truth when the breakage happens. The method allows us to establish the link betwixt a certain determinant through which maybe forecasted rock breakage and the talent to resist; the design can be grown and applied massif scale and again in the stability study of underground everything.

Author(s) Details:

Mihaela Toderas,
University of Petrosani, Petrosani-332006, Romania.

Please see the link here: https://stm.bookpi.org/ACST-V2/article/view/11930

Tuesday, 29 November 2022

A Handbook on Healthcare Applications | Book Publisher International

 With the rise in great dossier, machine intelligence has enhance specifically main for answering questions. Machine learning uses two types of techniques: directed education and alone education. Clustering is ultimate prevailing alone education method. Classification and Regression are directed education methods. Clustering algorithms attempt two broad groups: Hard grouping and soft grouping. K-Means, K-Mediods, Hierarchical grouping, Self-arranging Map are few of the hard assembling arrangements. Fuzzy C- Means, Gaussian Mixture Model are gentle grouping forms. In categorization question, the classes grant permission be twofold or multiclass. A multiclass categorization problem is mainly challenging cause it demands a more intricate model. Most ordinary categorization algorithms involve Logistic Regression, k Nearest Neighbor (kNN), Support Vector Machine (SVM), Neural Network, Naïve Bayes, Discriminant Analysis, Decision Tree, Bagged and Boosted Decision Trees. Regression algorithms contain Gaussian Process Regression Model, SVM Regression, Generalized Linear Model and Regression Tree.Depends on the request, few questions demand pre-handle and addition. Real-world datasets maybe dirty, unfinished and in a difference of layouts. Hence Pre-treat should before resolving the question. Machine learning is an persuasive system for verdict patterns in large datasets. But more generous data influences additional complicatedness. As datasets become larger, it is owned by decrease the number of facial characteristics. The three most usually secondhand range decline methods are: Principal Component Analysis (PCA), Factor Analysis and Nonnegative form factorization. The acting of the method seemingly increases when machine intelligence algorithms is secondhand. Selecting a machine intelligence treasure is a process of experimental approach. The particular traits of the algorithms contain Speed of preparation, Memory custom, Predictive veracity on new dossier, Transparency or interpretability.

Author(s) Details:

S. Sowmyayani,
Department of Computer Science (SF), St. Mary’s College (Autonomous), Thoothukudi, Tamilnadu, India.

Please see the link here: https://stm.bookpi.org/AHHA/article/view/8699

Monday, 17 January 2022

Global Impacts of Corruption Perception Index for the Attractiveness of Foreign Direct Investment | Chapter 11 | New Innovations in Economics, Business and Management Vol. 4

 The correlation between the perception of corruption and the attractiveness of foreign direct investment in the formation and implementation of state investment policy, as well as the impact of development projects in countries that use multiple regression analytical formulas, is investigated in this article. We also recognise several key drivers and determinants in modelling the issues of foreign direct investment, which are linked to attracting foreign direct investment and boosting the attractiveness of the economy's development. The investigation's major focus is on interpreters of econometric modelling of the relationship between selected nations' levels of corruption and the attractiveness of foreign direct investment. The real-life examples are related to corruption and foreign direct investment, both of which have been researched by various scientists around the world. The aim is to figure out how much corruption in the world affects the attractiveness of certain countries to international investors.


Author(S) Details

Ikboljon Odashev Mashrabjonovich
Institute for Forecasting and Macroeconomic Research, The Ministry of Economic Development and Poverty Reduction, Republic of Uzbekistan.

View Book:- https://stm.bookpi.org/NIEBM-V4/article/view/5338

Thursday, 26 August 2021

Study of the Compressive Strength Characteristics of Cement Mortar Reinforced with Kenaf Fibre | Chapter 14 | Recent Trends in Chemical and Material Sciences Vol. 2

 Kenaf fibre is one of the various products derived from the kenaf plant, which is found in abundance in many places of the globe. The use of kenaf fibre to strengthen cement mortars as a building material yielded positive results in research. Cement mortar mixtures were proportioned to include 1-3 percent fibre volume and 10-30mm fibre length. The composite's performance is measured in terms of density, water absorption, and compressive strength. According to the findings, water absorption increased with fibre volume but remained within ASTM C 211-77 limitations. The density of the kenaf fibre cement mortar did not change significantly. There is no statistical difference between the means of the compressive strength of plain mortar and those containing 1-3 percent fibre volume at 10mm fibre length, according to an analysis of variance (ANOVA) at a 5% level of significance. The compressive strength data revealed a substantial link between fibre volume, fibre length, and curing age when regression models were built.


Author (S) Details

T. M. Omoniyi
Civil Engineering Department, Abubakar Tafawa Balewa University, PMB 0248 Bauchi, Bauchi State, Nigeria.

Duna Samson
DG/CEO, Nigerian Building and Road Research Institute, Abuja, Nigeria.

Othman Musa Waila
Civil Engineering Department, Abubakar Tafawa Balewa University, PMB 0248 Bauchi, Bauchi State, Nigeria.

View Book :- https://stm.bookpi.org/RTCAMS-V2/article/view/2915

Wednesday, 30 June 2021

Soil Management Investment on Cassava Production in Oyo State, Nigeria | Chapter 6 | Recent Progress in Plant and Soil Research Vol. 1

 Using cross-sectional data, this study examines soil management investment in cassava production in Ido Local Government Area of Oyo State (Nigeria). Agriculture is a major economic sector in Nigeria, employing the vast majority of the population. Commercialization at the small, medium, and large scale enterprise levels is transforming the sector. Data were collected from 88 respondents using a structured questionnaire; four villages were chosen at random for the study. The collected data was analyzed using descriptive, mean, and multiple regression techniques. According to the findings, 84.1 percent of farmers were male, while 15.9 percent were female. 45.4 percent were between the ages of 21 and 30 years, 60.2 percent had 1-10 years of farming experience, and 33.0 percent had tertiary education. The most common soil management practices used by respondents were fertilizer and manure applications; 44.3 percent of farmers spent between N11,000 and N20,000 on soil management during the farming season. The regression analyses revealed that farm size and cassava output were both positively significant (= 0.203, p0.10) and (= 0.262, p0.01)1, respectively, whereas labor used was negatively significant (= -0.163, p0.01).  p0.10) to the level of soil management investment It was suggested, however, that farmers be better educated on appropriate soil management coping strategies. As a result, farmers should be encouraged by the government to improve their soil management system in order to increase productivity in the study area by providing formal credit facilities with no or low interest rates. It is suggested that policymakers assist farmers by providing agricultural credit and farm machinery at a subsidised rate, which could solve the problem of labor intensity and increase farmer productivity.

Author(s) Details

Isaac. O. Oyewo,
Federal College of Forestry (FRIN), P.M.B. 5087, Jericho Hill, Ibadan, Oyo State, Nigeria.

Dr. Adejare. A. Adesope
Forestry Research Institute of Nigeria, P.M.B. 5054, Jericho, Ibadan, Nigeria.

Esther. O. O. Ladipupo-Alade
Forestry Research Institute of Nigeria, P.M.B. 5054, Jericho, Ibadan, Nigeria.

View Book :- https://stm.bookpi.org/RPPSR-V1/article/view/1864