<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>http://practicalstats.labanca.net/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Kuslis001</id>
	<title>Practical Statistics for Educators - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="http://practicalstats.labanca.net/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Kuslis001"/>
	<link rel="alternate" type="text/html" href="http://practicalstats.labanca.net/index.php/Special:Contributions/Kuslis001"/>
	<updated>2026-09-25T01:13:11Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.31.16</generator>
	<entry>
		<id>http://practicalstats.labanca.net/index.php?title=Data_Screening&amp;diff=174</id>
		<title>Data Screening</title>
		<link rel="alternate" type="text/html" href="http://practicalstats.labanca.net/index.php?title=Data_Screening&amp;diff=174"/>
		<updated>2019-11-17T14:59:35Z</updated>

		<summary type="html">&lt;p&gt;Kuslis001: /* Detection of Multivariate Outliers: Scatterplot Matrices */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Data Screening ==&lt;br /&gt;
&lt;br /&gt;
Once data from a research study is gathered and has been entered into SPSS, researchers must examine their data to be sure they can validly interpret their results. Valid interpretation of data is reliant on two data features:&lt;br /&gt;
&lt;br /&gt;
1. The data must meet the assumptions of the analysis procedure.&lt;br /&gt;
&lt;br /&gt;
2. The data in the data file are &amp;quot;an accurate representation or transcription of what was provided by research participants as their original responses or what was provided by archival sources as original data&amp;quot; (Meyers, Gamst, &amp;amp; Guarino, 2017, p. 31).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;Contribution by: Britany Kuslis, WCSU Cohort 8&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
&lt;br /&gt;
Meyers, L., Gamst, G., &amp;amp; Guarino, A.J. (2017). &amp;#039;&amp;#039;Applied multivariate research: Design and interpretation.&amp;#039;&amp;#039; Thousand Oaks, CA: Sage Publications.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Data Cleaning ==&lt;br /&gt;
&lt;br /&gt;
This process of representing original data. In its initial phases, involves the researcher looking for incomplete data that may skew futher data screening. For example if there is a 5 question subscale where all the scores are added and a mean derived, if a respondent does not answer a question, that would skew the subscale, and that set of answers should be removed. &lt;br /&gt;
&lt;br /&gt;
Contribution by: Mykal Kuslis, WCSU Cohort 8&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
&lt;br /&gt;
Meyers, S., Gamst, G., &amp;amp; Guarino, A.J. (2017). Applied multivariate research: Design and interpretation. Thousand Oaks, CA: Sage Publications. (p. 31-33)&lt;br /&gt;
&lt;br /&gt;
== Value Cleaning ==&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Value cleaning&amp;#039;&amp;#039;&amp;#039; is ensuring the values are &amp;quot;within the limits of reasonable expectation&amp;quot; within the &amp;quot;to the extent that it is possible...within the bounds of feasibility&amp;quot;(Meyers, Gamst, &amp;amp; Guarino, 2017, p. 32). For example, you want to ensure the age of a presumed adult is not 9 years old or that a response to an item rated on a likert scale of 1-5 is not a 6 or an otherwise value that is not within the bounds of the study.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;Contribution by: Britany Kuslis, WCSU Cohort 8&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
Meyers, S., Gamst, G, &amp;amp; Guarino, A.J. (2017). &amp;#039;&amp;#039;Applied multivariate research: Design and interpretation.&amp;#039;&amp;#039; Thousand Oaks, CA: Sage Publications.&lt;br /&gt;
&lt;br /&gt;
== Outliers ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Outliers&amp;#039;&amp;#039;&amp;#039; are values that are &amp;quot;extreme or unusual values on a single variable (univariate) or on a combination of variables (multivariate)&amp;quot; (Meyers, Gamst, &amp;amp; Guarino, 2017, p. 48). &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The presence of outliers can greatly impact the results of an analysis for two major reasons: &lt;br /&gt;
&lt;br /&gt;
(1) The mean of the variable might no longer be a good variable and&lt;br /&gt;
 &lt;br /&gt;
(2) Outliers will yield a difference that when squared will produce a value too large that will skew the computation.  &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Outliers may signal &amp;quot;anomalies within the data&amp;quot; that will likely need to be addressed prior to moving forward with the statistical analysis (Meyers, Gamst, &amp;amp; Guarino, 2017, p. 48).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;Contribution by: Britany Kuslis, WCSU Cohort 8&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
&lt;br /&gt;
Meyers, S., Gamst, G., &amp;amp; Guarino, A.J. (2017). &amp;#039;&amp;#039;Applied multivariate research: Design and interpretation.&amp;#039;&amp;#039; Thousand Oaks, CA: Sage Publications.&lt;br /&gt;
&lt;br /&gt;
== Causes of Outliers ==&lt;br /&gt;
&lt;br /&gt;
- Data entry errors or improper attribute coding (normally caught in data cleaning)&lt;br /&gt;
&lt;br /&gt;
- A function of extraordinary events or unusual circumstances (for example a traumatic event causing someone to forget what they learned or a person remembering all 80 facts)&lt;br /&gt;
&lt;br /&gt;
- Some have no explanation, these are good cause for deletion.&lt;br /&gt;
&lt;br /&gt;
-Multivariate outliers- a pattern of combination of valuable on several variables.  &lt;br /&gt;
&lt;br /&gt;
Contribution by: Mykal Kuslis, WCSU Cohort 8&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
&lt;br /&gt;
Meyers, S., Gamst, G., &amp;amp; Guarino, A.J. (2017). Applied multivariate research: Design and interpretation. Thousand Oaks, CA: Sage Publications. (p.48-49)&lt;br /&gt;
&lt;br /&gt;
== Detection of Multivariate Outliers: Scatterplot Matrices ==&lt;br /&gt;
&lt;br /&gt;
Multivariate outliers uniqueness occurs in their pattern of combination of values on several variables. For example, a particular combination of age, sex, and number of arrests may be quite different from other combinations (young males in certain populations will have proportionally more arrests than other combinations of sex and age).&lt;br /&gt;
&lt;br /&gt;
Running bivariate scatterplots for combinations of key variables.&lt;br /&gt;
&lt;br /&gt;
Run a Scatterplot Matrices.&lt;br /&gt;
&lt;br /&gt;
Each case is represented as a point on the X and Y axes. &lt;br /&gt;
&lt;br /&gt;
Most cases will fall within the elliptical swarm or pattern mass, outliers are those cases that tend to lie outside the oval.&lt;br /&gt;
&lt;br /&gt;
See page 52 on reference for example matrices. &lt;br /&gt;
&lt;br /&gt;
Contribution by: Mykal Kuslis, WCSU Cohort 8&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
&lt;br /&gt;
Meyers, S., Gamst, G., &amp;amp; Guarino, A.J. (2017). Applied multivariate research: Design and interpretation. Thousand Oaks, CA: Sage Publications. (p.49-53)&lt;br /&gt;
&lt;br /&gt;
== Detection of Multivariate Outliers: Mahalanobis Distance ==&lt;br /&gt;
&lt;br /&gt;
The &amp;#039;&amp;#039;&amp;#039;Mahalanobis Distance&amp;#039;&amp;#039;&amp;#039; statistic measures &amp;quot;the multivariate &amp;#039;distance&amp;#039; between each case and the group multivariate mean (known as centroid) taking into account the correlations between the variables&amp;quot; (Meyers, Gamst, &amp;amp; Guarino, 2017, p. 52). This method is used to determine if there are scores that vary from the mean of a set of DV&amp;#039;s. The Mahalanobis distances details how far a case is from the group center mass of the predictor or IV&amp;#039;s.  The greater the distance the higher the possibility of a multivariate outlier.  According to Lawrence S. Meyers, Glenn Gamst and A.J. Guarino, &amp;quot;Each case is evaluated using the chi square distribution with a stringent alpha level of .001.  Cases that reach this significance threshold can be considered multivariate outliers and possible candidates for elimination. This approach is also not without its critics (e.g., Wilcox, 2012) for alternative approaches to multivariate outlier detection&amp;quot; (Meyers, Gamst, &amp;amp; Guarino, 2017, p.53).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Identifying Multivariate Outliers with Mahalanobis Distance--&amp;gt;[https://www.youtube.com/watch?v=AXLAX6r5JgE]&lt;br /&gt;
&lt;br /&gt;
Mahalanobis Distance --&amp;gt;[https://www.youtube.com/watch?v=spNpfmWZBmg]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;Contribution by: Britany Kuslis, WCSU Cohort 8&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
Clapham, Matthew E. “Mahalanobis Distance.” YouTube, YouTube.com, 2016, www.youtube.com/watch?v=spNpfmWZBmg.&lt;br /&gt;
&lt;br /&gt;
Grande, Dr. Todd. “Identifying Multivariate Outliers with Mahalanobis Distance.” YouTube, YouTube.com, 2016, www.youtube.com/watch?v=AXLAX6r5JgE.&lt;br /&gt;
&lt;br /&gt;
Meyers, L., Gamst, G, &amp;amp; Guarino, A.J. (2017). &amp;#039;&amp;#039;Applied multivariate research: Design and interpretation.&amp;#039;&amp;#039; Thousand Oaks, CA: Sage Publications.&lt;/div&gt;</summary>
		<author><name>Kuslis001</name></author>
		
	</entry>
	<entry>
		<id>http://practicalstats.labanca.net/index.php?title=Data_Screening&amp;diff=173</id>
		<title>Data Screening</title>
		<link rel="alternate" type="text/html" href="http://practicalstats.labanca.net/index.php?title=Data_Screening&amp;diff=173"/>
		<updated>2019-11-17T14:52:14Z</updated>

		<summary type="html">&lt;p&gt;Kuslis001: /* Detection of Multivariate Outliers */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Data Screening ==&lt;br /&gt;
&lt;br /&gt;
Once data from a research study is gathered and has been entered into SPSS, researchers must examine their data to be sure they can validly interpret their results. Valid interpretation of data is reliant on two data features:&lt;br /&gt;
&lt;br /&gt;
1. The data must meet the assumptions of the analysis procedure.&lt;br /&gt;
&lt;br /&gt;
2. The data in the data file are &amp;quot;an accurate representation or transcription of what was provided by research participants as their original responses or what was provided by archival sources as original data&amp;quot; (Meyers, Gamst, &amp;amp; Guarino, 2017, p. 31).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;Contribution by: Britany Kuslis, WCSU Cohort 8&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
&lt;br /&gt;
Meyers, L., Gamst, G., &amp;amp; Guarino, A.J. (2017). &amp;#039;&amp;#039;Applied multivariate research: Design and interpretation.&amp;#039;&amp;#039; Thousand Oaks, CA: Sage Publications.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Data Cleaning ==&lt;br /&gt;
&lt;br /&gt;
This process of representing original data. In its initial phases, involves the researcher looking for incomplete data that may skew futher data screening. For example if there is a 5 question subscale where all the scores are added and a mean derived, if a respondent does not answer a question, that would skew the subscale, and that set of answers should be removed. &lt;br /&gt;
&lt;br /&gt;
Contribution by: Mykal Kuslis, WCSU Cohort 8&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
&lt;br /&gt;
Meyers, S., Gamst, G., &amp;amp; Guarino, A.J. (2017). Applied multivariate research: Design and interpretation. Thousand Oaks, CA: Sage Publications. (p. 31-33)&lt;br /&gt;
&lt;br /&gt;
== Value Cleaning ==&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Value cleaning&amp;#039;&amp;#039;&amp;#039; is ensuring the values are &amp;quot;within the limits of reasonable expectation&amp;quot; within the &amp;quot;to the extent that it is possible...within the bounds of feasibility&amp;quot;(Meyers, Gamst, &amp;amp; Guarino, 2017, p. 32). For example, you want to ensure the age of a presumed adult is not 9 years old or that a response to an item rated on a likert scale of 1-5 is not a 6 or an otherwise value that is not within the bounds of the study.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;Contribution by: Britany Kuslis, WCSU Cohort 8&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
Meyers, S., Gamst, G, &amp;amp; Guarino, A.J. (2017). &amp;#039;&amp;#039;Applied multivariate research: Design and interpretation.&amp;#039;&amp;#039; Thousand Oaks, CA: Sage Publications.&lt;br /&gt;
&lt;br /&gt;
== Outliers ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Outliers&amp;#039;&amp;#039;&amp;#039; are values that are &amp;quot;extreme or unusual values on a single variable (univariate) or on a combination of variables (multivariate)&amp;quot; (Meyers, Gamst, &amp;amp; Guarino, 2017, p. 48). &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The presence of outliers can greatly impact the results of an analysis for two major reasons: &lt;br /&gt;
&lt;br /&gt;
(1) The mean of the variable might no longer be a good variable and&lt;br /&gt;
 &lt;br /&gt;
(2) Outliers will yield a difference that when squared will produce a value too large that will skew the computation.  &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Outliers may signal &amp;quot;anomalies within the data&amp;quot; that will likely need to be addressed prior to moving forward with the statistical analysis (Meyers, Gamst, &amp;amp; Guarino, 2017, p. 48).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;Contribution by: Britany Kuslis, WCSU Cohort 8&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
&lt;br /&gt;
Meyers, S., Gamst, G., &amp;amp; Guarino, A.J. (2017). &amp;#039;&amp;#039;Applied multivariate research: Design and interpretation.&amp;#039;&amp;#039; Thousand Oaks, CA: Sage Publications.&lt;br /&gt;
&lt;br /&gt;
== Causes of Outliers ==&lt;br /&gt;
&lt;br /&gt;
- Data entry errors or improper attribute coding (normally caught in data cleaning)&lt;br /&gt;
&lt;br /&gt;
- A function of extraordinary events or unusual circumstances (for example a traumatic event causing someone to forget what they learned or a person remembering all 80 facts)&lt;br /&gt;
&lt;br /&gt;
- Some have no explanation, these are good cause for deletion.&lt;br /&gt;
&lt;br /&gt;
-Multivariate outliers- a pattern of combination of valuable on several variables.  &lt;br /&gt;
&lt;br /&gt;
Contribution by: Mykal Kuslis, WCSU Cohort 8&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
&lt;br /&gt;
Meyers, S., Gamst, G., &amp;amp; Guarino, A.J. (2017). Applied multivariate research: Design and interpretation. Thousand Oaks, CA: Sage Publications. (p.48-49)&lt;br /&gt;
&lt;br /&gt;
== Detection of Multivariate Outliers: Scatterplot Matrices ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Detection of Multivariate Outliers: Mahalanobis Distance ==&lt;br /&gt;
&lt;br /&gt;
The &amp;#039;&amp;#039;&amp;#039;Mahalanobis Distance&amp;#039;&amp;#039;&amp;#039; statistic measures &amp;quot;the multivariate &amp;#039;distance&amp;#039; between each case and the group multivariate mean (known as centroid) taking into account the correlations between the variables&amp;quot; (Meyers, Gamst, &amp;amp; Guarino, 2017, p. 52). This method is used to determine if there are scores that vary from the mean of a set of DV&amp;#039;s. The Mahalanobis distances details how far a case is from the group center mass of the predictor or IV&amp;#039;s.  The greater the distance the higher the possibility of a multivariate outlier.  According to Lawrence S. Meyers, Glenn Gamst and A.J. Guarino, &amp;quot;Each case is evaluated using the chi square distribution with a stringent alpha level of .001.  Cases that reach this significance threshold can be considered multivariate outliers and possible candidates for elimination. This approach is also not without its critics (e.g., Wilcox, 2012) for alternative approaches to multivariate outlier detection&amp;quot; (Meyers, Gamst, &amp;amp; Guarino, 2017, p.53).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Identifying Multivariate Outliers with Mahalanobis Distance--&amp;gt;[https://www.youtube.com/watch?v=AXLAX6r5JgE]&lt;br /&gt;
&lt;br /&gt;
Mahalanobis Distance --&amp;gt;[https://www.youtube.com/watch?v=spNpfmWZBmg]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;Contribution by: Britany Kuslis, WCSU Cohort 8&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
Clapham, Matthew E. “Mahalanobis Distance.” YouTube, YouTube.com, 2016, www.youtube.com/watch?v=spNpfmWZBmg.&lt;br /&gt;
&lt;br /&gt;
Grande, Dr. Todd. “Identifying Multivariate Outliers with Mahalanobis Distance.” YouTube, YouTube.com, 2016, www.youtube.com/watch?v=AXLAX6r5JgE.&lt;br /&gt;
&lt;br /&gt;
Meyers, L., Gamst, G, &amp;amp; Guarino, A.J. (2017). &amp;#039;&amp;#039;Applied multivariate research: Design and interpretation.&amp;#039;&amp;#039; Thousand Oaks, CA: Sage Publications.&lt;/div&gt;</summary>
		<author><name>Kuslis001</name></author>
		
	</entry>
	<entry>
		<id>http://practicalstats.labanca.net/index.php?title=Data_Screening&amp;diff=172</id>
		<title>Data Screening</title>
		<link rel="alternate" type="text/html" href="http://practicalstats.labanca.net/index.php?title=Data_Screening&amp;diff=172"/>
		<updated>2019-11-17T14:50:56Z</updated>

		<summary type="html">&lt;p&gt;Kuslis001: /* Causes of Outliers */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Data Screening ==&lt;br /&gt;
&lt;br /&gt;
Once data from a research study is gathered and has been entered into SPSS, researchers must examine their data to be sure they can validly interpret their results. Valid interpretation of data is reliant on two data features:&lt;br /&gt;
&lt;br /&gt;
1. The data must meet the assumptions of the analysis procedure.&lt;br /&gt;
&lt;br /&gt;
2. The data in the data file are &amp;quot;an accurate representation or transcription of what was provided by research participants as their original responses or what was provided by archival sources as original data&amp;quot; (Meyers, Gamst, &amp;amp; Guarino, 2017, p. 31).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;Contribution by: Britany Kuslis, WCSU Cohort 8&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
&lt;br /&gt;
Meyers, L., Gamst, G., &amp;amp; Guarino, A.J. (2017). &amp;#039;&amp;#039;Applied multivariate research: Design and interpretation.&amp;#039;&amp;#039; Thousand Oaks, CA: Sage Publications.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Data Cleaning ==&lt;br /&gt;
&lt;br /&gt;
This process of representing original data. In its initial phases, involves the researcher looking for incomplete data that may skew futher data screening. For example if there is a 5 question subscale where all the scores are added and a mean derived, if a respondent does not answer a question, that would skew the subscale, and that set of answers should be removed. &lt;br /&gt;
&lt;br /&gt;
Contribution by: Mykal Kuslis, WCSU Cohort 8&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
&lt;br /&gt;
Meyers, S., Gamst, G., &amp;amp; Guarino, A.J. (2017). Applied multivariate research: Design and interpretation. Thousand Oaks, CA: Sage Publications. (p. 31-33)&lt;br /&gt;
&lt;br /&gt;
== Value Cleaning ==&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Value cleaning&amp;#039;&amp;#039;&amp;#039; is ensuring the values are &amp;quot;within the limits of reasonable expectation&amp;quot; within the &amp;quot;to the extent that it is possible...within the bounds of feasibility&amp;quot;(Meyers, Gamst, &amp;amp; Guarino, 2017, p. 32). For example, you want to ensure the age of a presumed adult is not 9 years old or that a response to an item rated on a likert scale of 1-5 is not a 6 or an otherwise value that is not within the bounds of the study.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;Contribution by: Britany Kuslis, WCSU Cohort 8&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
Meyers, S., Gamst, G, &amp;amp; Guarino, A.J. (2017). &amp;#039;&amp;#039;Applied multivariate research: Design and interpretation.&amp;#039;&amp;#039; Thousand Oaks, CA: Sage Publications.&lt;br /&gt;
&lt;br /&gt;
== Outliers ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Outliers&amp;#039;&amp;#039;&amp;#039; are values that are &amp;quot;extreme or unusual values on a single variable (univariate) or on a combination of variables (multivariate)&amp;quot; (Meyers, Gamst, &amp;amp; Guarino, 2017, p. 48). &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The presence of outliers can greatly impact the results of an analysis for two major reasons: &lt;br /&gt;
&lt;br /&gt;
(1) The mean of the variable might no longer be a good variable and&lt;br /&gt;
 &lt;br /&gt;
(2) Outliers will yield a difference that when squared will produce a value too large that will skew the computation.  &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Outliers may signal &amp;quot;anomalies within the data&amp;quot; that will likely need to be addressed prior to moving forward with the statistical analysis (Meyers, Gamst, &amp;amp; Guarino, 2017, p. 48).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;Contribution by: Britany Kuslis, WCSU Cohort 8&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
&lt;br /&gt;
Meyers, S., Gamst, G., &amp;amp; Guarino, A.J. (2017). &amp;#039;&amp;#039;Applied multivariate research: Design and interpretation.&amp;#039;&amp;#039; Thousand Oaks, CA: Sage Publications.&lt;br /&gt;
&lt;br /&gt;
== Causes of Outliers ==&lt;br /&gt;
&lt;br /&gt;
- Data entry errors or improper attribute coding (normally caught in data cleaning)&lt;br /&gt;
&lt;br /&gt;
- A function of extraordinary events or unusual circumstances (for example a traumatic event causing someone to forget what they learned or a person remembering all 80 facts)&lt;br /&gt;
&lt;br /&gt;
- Some have no explanation, these are good cause for deletion.&lt;br /&gt;
&lt;br /&gt;
-Multivariate outliers- a pattern of combination of valuable on several variables.  &lt;br /&gt;
&lt;br /&gt;
Contribution by: Mykal Kuslis, WCSU Cohort 8&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
&lt;br /&gt;
Meyers, S., Gamst, G., &amp;amp; Guarino, A.J. (2017). Applied multivariate research: Design and interpretation. Thousand Oaks, CA: Sage Publications. (p.48-49)&lt;br /&gt;
&lt;br /&gt;
== Detection of Multivariate Outliers ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Detection of Multivariate Outliers: Scatterplot Matrices ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Detection of Multivariate Outliers: Mahalanobis Distance ==&lt;br /&gt;
&lt;br /&gt;
The &amp;#039;&amp;#039;&amp;#039;Mahalanobis Distance&amp;#039;&amp;#039;&amp;#039; statistic measures &amp;quot;the multivariate &amp;#039;distance&amp;#039; between each case and the group multivariate mean (known as centroid) taking into account the correlations between the variables&amp;quot; (Meyers, Gamst, &amp;amp; Guarino, 2017, p. 52). This method is used to determine if there are scores that vary from the mean of a set of DV&amp;#039;s. The Mahalanobis distances details how far a case is from the group center mass of the predictor or IV&amp;#039;s.  The greater the distance the higher the possibility of a multivariate outlier.  According to Lawrence S. Meyers, Glenn Gamst and A.J. Guarino, &amp;quot;Each case is evaluated using the chi square distribution with a stringent alpha level of .001.  Cases that reach this significance threshold can be considered multivariate outliers and possible candidates for elimination. This approach is also not without its critics (e.g., Wilcox, 2012) for alternative approaches to multivariate outlier detection&amp;quot; (Meyers, Gamst, &amp;amp; Guarino, 2017, p.53).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Identifying Multivariate Outliers with Mahalanobis Distance--&amp;gt;[https://www.youtube.com/watch?v=AXLAX6r5JgE]&lt;br /&gt;
&lt;br /&gt;
Mahalanobis Distance --&amp;gt;[https://www.youtube.com/watch?v=spNpfmWZBmg]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;Contribution by: Britany Kuslis, WCSU Cohort 8&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
Clapham, Matthew E. “Mahalanobis Distance.” YouTube, YouTube.com, 2016, www.youtube.com/watch?v=spNpfmWZBmg.&lt;br /&gt;
&lt;br /&gt;
Grande, Dr. Todd. “Identifying Multivariate Outliers with Mahalanobis Distance.” YouTube, YouTube.com, 2016, www.youtube.com/watch?v=AXLAX6r5JgE.&lt;br /&gt;
&lt;br /&gt;
Meyers, L., Gamst, G, &amp;amp; Guarino, A.J. (2017). &amp;#039;&amp;#039;Applied multivariate research: Design and interpretation.&amp;#039;&amp;#039; Thousand Oaks, CA: Sage Publications.&lt;/div&gt;</summary>
		<author><name>Kuslis001</name></author>
		
	</entry>
	<entry>
		<id>http://practicalstats.labanca.net/index.php?title=Data_Screening&amp;diff=171</id>
		<title>Data Screening</title>
		<link rel="alternate" type="text/html" href="http://practicalstats.labanca.net/index.php?title=Data_Screening&amp;diff=171"/>
		<updated>2019-11-17T14:41:50Z</updated>

		<summary type="html">&lt;p&gt;Kuslis001: /* Data Cleaning */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Data Screening ==&lt;br /&gt;
&lt;br /&gt;
Once data from a research study is gathered and has been entered into SPSS, researchers must examine their data to be sure they can validly interpret their results. Valid interpretation of data is reliant on two data features:&lt;br /&gt;
&lt;br /&gt;
1. The data must meet the assumptions of the analysis procedure.&lt;br /&gt;
&lt;br /&gt;
2. The data in the data file are &amp;quot;an accurate representation or transcription of what was provided by research participants as their original responses or what was provided by archival sources as original data&amp;quot; (Meyers, Gamst, &amp;amp; Guarino, 2017, p. 31).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;Contribution by: Britany Kuslis, WCSU Cohort 8&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
&lt;br /&gt;
Meyers, L., Gamst, G., &amp;amp; Guarino, A.J. (2017). &amp;#039;&amp;#039;Applied multivariate research: Design and interpretation.&amp;#039;&amp;#039; Thousand Oaks, CA: Sage Publications.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Data Cleaning ==&lt;br /&gt;
&lt;br /&gt;
This process of representing original data. In its initial phases, involves the researcher looking for incomplete data that may skew futher data screening. For example if there is a 5 question subscale where all the scores are added and a mean derived, if a respondent does not answer a question, that would skew the subscale, and that set of answers should be removed. &lt;br /&gt;
&lt;br /&gt;
Contribution by: Mykal Kuslis, WCSU Cohort 8&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
&lt;br /&gt;
Meyers, S., Gamst, G., &amp;amp; Guarino, A.J. (2017). Applied multivariate research: Design and interpretation. Thousand Oaks, CA: Sage Publications. (p. 31-33)&lt;br /&gt;
&lt;br /&gt;
== Value Cleaning ==&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Value cleaning&amp;#039;&amp;#039;&amp;#039; is ensuring the values are &amp;quot;within the limits of reasonable expectation&amp;quot; within the &amp;quot;to the extent that it is possible...within the bounds of feasibility&amp;quot;(Meyers, Gamst, &amp;amp; Guarino, 2017, p. 32). For example, you want to ensure the age of a presumed adult is not 9 years old or that a response to an item rated on a likert scale of 1-5 is not a 6 or an otherwise value that is not within the bounds of the study.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;Contribution by: Britany Kuslis, WCSU Cohort 8&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
Meyers, S., Gamst, G, &amp;amp; Guarino, A.J. (2017). &amp;#039;&amp;#039;Applied multivariate research: Design and interpretation.&amp;#039;&amp;#039; Thousand Oaks, CA: Sage Publications.&lt;br /&gt;
&lt;br /&gt;
== Outliers ==&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Outliers&amp;#039;&amp;#039;&amp;#039; are values that are &amp;quot;extreme or unusual values on a single variable (univariate) or on a combination of variables (multivariate)&amp;quot; (Meyers, Gamst, &amp;amp; Guarino, 2017, p. 48). &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The presence of outliers can greatly impact the results of an analysis for two major reasons: &lt;br /&gt;
&lt;br /&gt;
(1) The mean of the variable might no longer be a good variable and&lt;br /&gt;
 &lt;br /&gt;
(2) Outliers will yield a difference that when squared will produce a value too large that will skew the computation.  &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Outliers may signal &amp;quot;anomalies within the data&amp;quot; that will likely need to be addressed prior to moving forward with the statistical analysis (Meyers, Gamst, &amp;amp; Guarino, 2017, p. 48).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;Contribution by: Britany Kuslis, WCSU Cohort 8&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
&lt;br /&gt;
Meyers, S., Gamst, G., &amp;amp; Guarino, A.J. (2017). &amp;#039;&amp;#039;Applied multivariate research: Design and interpretation.&amp;#039;&amp;#039; Thousand Oaks, CA: Sage Publications.&lt;br /&gt;
&lt;br /&gt;
== Causes of Outliers ==&lt;br /&gt;
&lt;br /&gt;
== Detection of Multivariate Outliers ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Detection of Multivariate Outliers: Scatterplot Matrices ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Detection of Multivariate Outliers: Mahalanobis Distance ==&lt;br /&gt;
&lt;br /&gt;
The &amp;#039;&amp;#039;&amp;#039;Mahalanobis Distance&amp;#039;&amp;#039;&amp;#039; statistic measures &amp;quot;the multivariate &amp;#039;distance&amp;#039; between each case and the group multivariate mean (known as centroid) taking into account the correlations between the variables&amp;quot; (Meyers, Gamst, &amp;amp; Guarino, 2017, p. 52). This method is used to determine if there are scores that vary from the mean of a set of DV&amp;#039;s. The Mahalanobis distances details how far a case is from the group center mass of the predictor or IV&amp;#039;s.  The greater the distance the higher the possibility of a multivariate outlier.  According to Lawrence S. Meyers, Glenn Gamst and A.J. Guarino, &amp;quot;Each case is evaluated using the chi square distribution with a stringent alpha level of .001.  Cases that reach this significance threshold can be considered multivariate outliers and possible candidates for elimination. This approach is also not without its critics (e.g., Wilcox, 2012) for alternative approaches to multivariate outlier detection&amp;quot; (Meyers, Gamst, &amp;amp; Guarino, 2017, p.53).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Identifying Multivariate Outliers with Mahalanobis Distance--&amp;gt;[https://www.youtube.com/watch?v=AXLAX6r5JgE]&lt;br /&gt;
&lt;br /&gt;
Mahalanobis Distance --&amp;gt;[https://www.youtube.com/watch?v=spNpfmWZBmg]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;Contribution by: Britany Kuslis, WCSU Cohort 8&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
Clapham, Matthew E. “Mahalanobis Distance.” YouTube, YouTube.com, 2016, www.youtube.com/watch?v=spNpfmWZBmg.&lt;br /&gt;
&lt;br /&gt;
Grande, Dr. Todd. “Identifying Multivariate Outliers with Mahalanobis Distance.” YouTube, YouTube.com, 2016, www.youtube.com/watch?v=AXLAX6r5JgE.&lt;br /&gt;
&lt;br /&gt;
Meyers, L., Gamst, G, &amp;amp; Guarino, A.J. (2017). &amp;#039;&amp;#039;Applied multivariate research: Design and interpretation.&amp;#039;&amp;#039; Thousand Oaks, CA: Sage Publications.&lt;/div&gt;</summary>
		<author><name>Kuslis001</name></author>
		
	</entry>
	<entry>
		<id>http://practicalstats.labanca.net/index.php?title=Confidence_Intervals&amp;diff=170</id>
		<title>Confidence Intervals</title>
		<link rel="alternate" type="text/html" href="http://practicalstats.labanca.net/index.php?title=Confidence_Intervals&amp;diff=170"/>
		<updated>2019-11-17T14:26:51Z</updated>

		<summary type="html">&lt;p&gt;Kuslis001: Created page with &amp;quot; == Creating Confidence Intervals ==  The use of confidence intervals is in part, due to the fact that the traditional and restricted framework of statistical significance tes...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Creating Confidence Intervals ==&lt;br /&gt;
&lt;br /&gt;
The use of confidence intervals is in part, due to the fact that the traditional and restricted framework of statistical significance testing has not been universally endorsed, therefore creating the need for confidence intervals.&lt;br /&gt;
&lt;br /&gt;
This comes down to a simple question, &amp;quot;Is it possible to assert something positive and tangible about the means of the groups in an experimental study?&amp;quot;&lt;br /&gt;
&lt;br /&gt;
Instead of using significance level in a study, it maybe more beneficial to use a confidence interval (which is the opposite of the significance level).&lt;br /&gt;
&lt;br /&gt;
For example, saying &amp;quot;the 6 month survival rate wan increased by 30 percentage points with a 99% confidence interval&amp;quot; than by simple saying the difference between the control group and experimental group was significant at the .01 level.&lt;br /&gt;
&lt;br /&gt;
The creation of the confidence interval then, becomes the percentage remaining from the significance level. In this this case 100-1= 99%&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Contribution by: Mykal Kuslis, WCSU Cohort 8&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
&lt;br /&gt;
Meyers, S., Gamst, G., &amp;amp; Guarino, A.J. (2017). Applied multivariate research: Design and interpretation. Thousand Oaks, CA: Sage Publications. (p.24-25)&lt;/div&gt;</summary>
		<author><name>Kuslis001</name></author>
		
	</entry>
	<entry>
		<id>http://practicalstats.labanca.net/index.php?title=Pearson_r&amp;diff=169</id>
		<title>Pearson r</title>
		<link rel="alternate" type="text/html" href="http://practicalstats.labanca.net/index.php?title=Pearson_r&amp;diff=169"/>
		<updated>2019-11-17T14:17:22Z</updated>

		<summary type="html">&lt;p&gt;Kuslis001: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Also known as Pearson&amp;#039;s product-moment correlation.  This technique is used to correlate the raw scores of two variables.&lt;br /&gt;
&lt;br /&gt;
Also visit http://psych.csufresno.edu/psy144/Content/Statistics/relationship_strength.html for more information on Pearson r.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;contributed by Kara Kunst&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also referred to as the Pearson Correlation Coefficient Squared, it is the proportion of variance in the criterion variable that can be accounted for by the predictor variable. (from Dr. Nancy Heilbronner)&lt;br /&gt;
&lt;br /&gt;
&amp;quot;contributed by Mary Fernand&amp;quot;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
Note: Pearson r scores cannot exceed 1.00 or -1.00 (range is between -1.00 and 1.00). &lt;br /&gt;
&lt;br /&gt;
The Pearson r score (say for example .80) is the number where the distribution will peak, and the remaining distribution will spread out around the number. &lt;br /&gt;
&lt;br /&gt;
Contribution by: Mykal Kuslis, WCSU Cohort 8&lt;br /&gt;
&lt;br /&gt;
Reference:&lt;br /&gt;
&lt;br /&gt;
Meyers, S., Gamst, G., &amp;amp; Guarino, A.J. (2017). Applied multivariate research: Design and interpretation. Thousand Oaks, CA: Sage Publications. (P. 21)&lt;/div&gt;</summary>
		<author><name>Kuslis001</name></author>
		
	</entry>
	<entry>
		<id>http://practicalstats.labanca.net/index.php?title=Pearson_r&amp;diff=168</id>
		<title>Pearson r</title>
		<link rel="alternate" type="text/html" href="http://practicalstats.labanca.net/index.php?title=Pearson_r&amp;diff=168"/>
		<updated>2019-11-17T14:09:37Z</updated>

		<summary type="html">&lt;p&gt;Kuslis001: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Also known as Pearson&amp;#039;s product-moment correlation.  This technique is used to correlate the raw scores of two variables.&lt;br /&gt;
&lt;br /&gt;
Also visit http://psych.csufresno.edu/psy144/Content/Statistics/relationship_strength.html for more information on Pearson r.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;contributed by Kara Kunst&amp;#039;&amp;#039;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Also referred to as the Pearson Correlation Coefficient Squared, it is the proportion of variance in the criterion variable that can be accounted for by the predictor variable. (from Dr. Nancy Heilbronner)&lt;br /&gt;
&lt;br /&gt;
&amp;quot;contributed by Mary Fernand&amp;quot;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
----&lt;/div&gt;</summary>
		<author><name>Kuslis001</name></author>
		
	</entry>
	<entry>
		<id>http://practicalstats.labanca.net/index.php?title=Contributions_here&amp;diff=167</id>
		<title>Contributions here</title>
		<link rel="alternate" type="text/html" href="http://practicalstats.labanca.net/index.php?title=Contributions_here&amp;diff=167"/>
		<updated>2019-11-17T14:03:09Z</updated>

		<summary type="html">&lt;p&gt;Kuslis001: /* Student Contributors */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Editor ==&lt;br /&gt;
Frank LaBanca, EdD&lt;br /&gt;
&lt;br /&gt;
== Faculty Contributors ==&lt;br /&gt;
Karen Burke, EdD&lt;br /&gt;
&lt;br /&gt;
Patricia Cosentino, EdD&lt;br /&gt;
&lt;br /&gt;
Deborah Hardy, EdD&lt;br /&gt;
&lt;br /&gt;
Jennifer Mitchell, EdD&lt;br /&gt;
&lt;br /&gt;
== Student Contributors ==&lt;br /&gt;
David Bozzuto&lt;br /&gt;
&lt;br /&gt;
Karen Fildes&lt;br /&gt;
&lt;br /&gt;
Michael Minzloff&lt;br /&gt;
&lt;br /&gt;
Damien Holst&lt;br /&gt;
&lt;br /&gt;
Jennifer Eraca&lt;br /&gt;
&lt;br /&gt;
John Ryan&lt;br /&gt;
&lt;br /&gt;
Kara Kunst&lt;br /&gt;
&lt;br /&gt;
Emily Rhew&lt;br /&gt;
&lt;br /&gt;
Cassandra Cosentino&lt;br /&gt;
&lt;br /&gt;
Kristina Hislop&lt;br /&gt;
&lt;br /&gt;
Mary Fernand&lt;br /&gt;
&lt;br /&gt;
Thomas Fox&lt;br /&gt;
&lt;br /&gt;
Helen Knudsen&lt;br /&gt;
&lt;br /&gt;
Ashley Brooksbank&lt;br /&gt;
&lt;br /&gt;
Scott Trungadi&lt;br /&gt;
&lt;br /&gt;
Sheri Prendergast&lt;br /&gt;
&lt;br /&gt;
Britany Kuslis&lt;br /&gt;
&lt;br /&gt;
Mykal Kuslis&lt;/div&gt;</summary>
		<author><name>Kuslis001</name></author>
		
	</entry>
</feed>