CSCS Study Guide: Chapter 13–Test Selection and Adminstration

CSCS Study Guides

CSCS Chapter 13 Study Guide

Test Selection and Administration

This is my Chapter 13 CSCS study guide for Test Selection and Administration. The big themes are choosing tests that match the sport, knowing the difference between validity and reliability, controlling the testing environment, and sequencing tests so fatigue does not ruin the data.

This guide is designed as a quick-reference review, not a substitute for the full textbook or hands-on testing practice.

How to use this guide
First pass: memorize the validity and reliability definitions.
Second pass: learn the testing order until it is automatic.
Third pass: quiz the trapboard. The CSCS goblin loves asking whether a bad test is invalid, unreliable, or just poorly administered.

1. High-Yield Snapshot

ConceptCSCS Exam Version
Main purpose of testingAssess ability and progress over time; identify needs; set reasonable objectives; evaluate program effectiveness.
TestAn assessment of ability.
Field testA test performed in a natural environment with minimal equipment or specialized training.
MeasurementCollection of testing data. Data collection = measurement.
EvaluationAnalyzing test results and using them to make decisions.
ReliabilityConsistency or repeatability of a test.
ValidityDegree to which a test measures what it claims to measure.
Testing orderNonfatiguing -> agility -> max power/strength -> sprint -> local muscular endurance -> anaerobic capacity -> aerobic capacity.
Rest between near-max attemptsAt least 3 minutes.
Rest between tests in a batteryAt least 5 minutes to limit fatigue carryover.
Glycolytic recovery after maximal anaerobic capacity test1 hour or more.
Altitude adjustmentAerobic norms may need adjustment above about 1,900 ft; VO2max drops roughly 5% per 3,000 ft gain.
Exam Spell
Measurement collects. Evaluation decides. Reliability repeats. Validity measures the right thing. Specificity makes the test look like the sport demand.

2. Why Test Athletes?

  • Talent identification: Determine whether an athlete has the physical potential to participate at a competitive level or in a specific position.
  • Needs analysis: Find performance areas that require improvement.
  • Program evaluation: Determine whether the training program is producing the desired adaptations.
  • Goal setting: Set realistic, measurable objectives.
  • Progress tracking: Compare pre-test, mid-test, and post-test results.
  • Risk awareness: Screen for concerns that may warrant referral or modified testing, while remembering that formal diagnosis is outside the strength coach role.
  • Not a valid reason: to fill a training-hour quota, entertain the team, or test for the sake of testing.
  • Testing should answer a coaching question. If the test does not change decisions, the test may be gym confetti.

3. Evidence-Based Testing and the Scientific Process

  • Evidence-based testing: Use peer-reviewed research, textbooks, professional organizations, meta-analyses, and field experience to choose protocols.
  • Scientific process: Ask a question, select a test, form a hypothesis, apply an intervention, evaluate results, and adjust.
StepVolleyball Example
QuestionHow can I evaluate middle-blocker potential?
TestWeekly approach jump testing.
HypothesisTracking approach jump helps identify or develop middle-blocker performance potential.
InterventionImplement appropriate plyometric training volume.
EvaluationReview results and either continue, modify the test, or adjust the intervention.

4. Timing of Tests and Evaluations

TermMeaningExam Cue
Pre-testBefore training begins; establishes baseline.Think: starting line.
Mid-test / Intra-training testOne or more tests during a training block.Purpose: evaluate progress and modify the program.
Formative evaluationOngoing/periodic evaluation using mid-test data.Progress monitoring and program adjustment.
Post-test / Summative evaluationAfter the training period.Determines whether the program achieved its goals.

5. Validity: Does the Test Measure the Right Thing?

  • A test can be reliable but not valid. It can consistently measure the wrong thing.
  • A test cannot be valid if it is not reliable, because inconsistent data cannot support meaningful interpretation.
  • Construct validity is the big umbrella: how well the test measures the construct it claims to measure.
Validity TypeDefinitionCSCS Example / Trap
Construct validityOverall degree that a test measures what it is designed to measure.Most important for overall test quality.
Face validityThe test appears valid to athletes or observers.Least rigorous. If it looks like it measures speed, athletes may believe it has face validity.
Content validityExperts judge whether the test covers all relevant subtopics/component abilities in proper proportion.Expert judgment = content validity, not criterion validity.
Criterion-referenced validityScores are associated with another measure of the same ability.Includes concurrent and predictive validity.
Concurrent validityScores are associated with an accepted test of the same ability at about the same time.Underwater weighing close to DEXA = high concurrent validity.
Predictive validityTest score is associated with future behavior or performance.Recruitment battery predicting future volleyball performance.
Convergent validityHigh positive correlation with a gold standard or closely related construct.New sprint test correlates with gold-standard timing.
Discriminant validityLow correlation with a different construct; can distinguish different abilities/groups.Vertical jump should not correlate too strongly with unrelated endurance ability.
Validity Trapboard
Gold standard comparison = convergent or concurrent depending wording.
Future performance = predictive.
Expert judgment of coverage = content.
Looks valid to athletes = face.
Different construct / low correlation = discriminant.

6. Reliability: Can the Test Repeat Cleanly?

  • Reliability is the consistency or repeatability of a test.
  • High intersubject variability does not automatically make a test unreliable; athletes can be different from each other. The issue is whether each athlete and each tester produces consistent results under standardized conditions.
Reliability / Error TermDefinitionExam Cue
Test-retest reliabilityStatistical correlation of scores from two administrations of the same test to the same group.Monday and Wednesday vertical jumps under identical conditions produce similar results.
Intraclass correlation coefficient (ICC)Statistic used to quantify test-retest reliability or agreement.Often appears as the test-retest statistic.
Intrasubject variabilityLack of consistent performance by the person being tested.Same athlete is inconsistent.
Interrater reliability / objectivityAgreement between different raters/testers.Two coaches timing the same pro-agility test should agree.
Intrarater variabilityLack of consistent scores from the same tester.Same coach scores differently across attempts/days.
Typical error of measurement (TE)Calculated statistic that includes equipment error, biological variation, and other measurement noise.The fuzz around the number.

7. Test Selection: Match the Test to the Question

  • Sport specificity is king: Choose tests that mimic the sport demands, not tests that are merely convenient.
  • Movement pattern specificity: Match the relevant movement pattern: vertical jump for basketball/volleyball, 15-yard sprint for an offensive lineman, modified horizontal pulling endurance for rowing.
  • Metabolic specificity: Match energy system demands: anaerobic tests for intermittent court sports, endurance tests for endurance athletes.
  • Muscle group specificity: Test the muscles that matter for the sport or position.
  • Skill level / training status: Technique-intensive tests are more appropriate for trained athletes; less trained athletes may need safer substitutes.
  • Age and sex considerations: Account for youth experience, maturation, sex-related strength differences, and test appropriateness.
  • Environment: Control heat, humidity, altitude, time of day, day of week, and season whenever possible.

8. Common Test Categories and What They Measure

CategoryExamplesWhat It Measures / Notes
Aerobic capacity1.5-mile run, 12-minute run, Yo-Yo testsCardiorespiratory endurance. Aerobic tests go last or on a separate day.
AgilityT-test, pro-agility/5-10-5, arrowhead testChange-of-direction ability. Agility is tested early, before strength and sprints.
Anaerobic capacity300-yard shuttle, Wingate testHigh-intensity capacity with glycolytic demand. Requires long recovery afterward.
Balance/stabilityBESS, Star Excursion Balance Test, Y-Balance TestPostural control and dynamic balance/stability.
Body compositionDXA, skinfolds, BIA, circumferences, underwater weighingFat and fat-free mass estimates. BMI is limited in athletic populations.
Maximum muscular powerVertical jump, standing long jump, Margaria-Kalamen testExplosive power. Vertical jump is nonfatiguing and can appear early in sequence.
Maximum muscular strength1RM bench/squat/power clean; 3RM or 10RM estimates when saferStrength. Technique and safety determine test choice.
Speed10-40 yard/meter sprint testsAcceleration or top-end speed depending distance.
Local muscular endurancePush-up test, YMCA bench press test, prone pull-up testRepeated submaximal output in a local muscle group.
Position-Specific Examples
Volleyball middle blocker: approach jump or vertical jump for jumping ability; 300-yard shuttle if the question asks anaerobic sport demands.
Defensive lineman: 15-yard sprint plus 1RM bench is more specific than long sprint or elbow-flexion endurance.
Cross-country runner: 1.5-mile run is more specific than T-test or Margaria-Kalamen.
Javelin thrower: height/body measurements first if included, then explosive tests later in the proper sequence.

9. Safety, Environment, and Referral Red Flags

  • The coach must observe for signs and symptoms that warrant exclusion from testing, especially during maximal exertion tests.
  • Medical referral should be made at the tester’s discretion when concerning signs appear.
  • Hot and humid environments can impair endurance performance, increase health risk, and lower validity of aerobic tests.
  • Environmental conditions should be documented and standardized when possible.
IssueSigns / Exam CluesAction / Study Note
Heat strokeCramps, nausea, red skin, absent sweating or altered mental status can indicate serious heat illness.Medical emergency. Stop testing and activate emergency response.
Heat exhaustion / heat illness riskHeavy sweating, weakness, dizziness, nausea, rapid pulse, fatigue.Stop/modify testing; cool athlete; refer when warranted.
HyponatremiaDilute urine, bloated skin, confusion, nausea, headache, swelling after excessive fluid intake.Electrolyte problem. Do not just push more plain water.
Cardiac warning signsChest pressure/pain/discomfort, fainting, unusual shortness of breath.Warrants medical attention.
AltitudeReduced aerobic performance due to lower oxygen availability.Adjust aerobic norms above ~1,900 ft; ~5% VO2max decline per 3,000 ft gain.
High body mass / body compositionBMI can misclassify muscular athletes.Use direct or composition-sensitive measures when possible.

10. Test Administration: Make the Data Trustworthy

  • Selection and training of testers: Testers must understand protocols, practice procedures, and give consistent explanations. This improves reliability, especially interrater reliability.
  • Recording forms: Prepare scoring forms ahead of time with spaces for results, comments, environmental conditions, and abnormal observations.
  • Test format: Decide whether testing is individual or group-based. The same tester should administer a given test whenever possible.
  • Warm-up: Use a standardized general warm-up followed by a specific warm-up. Warm-ups increase reliability.
  • Practice/familiarization: A low-exertion pre-test 1-3 days before the actual test can improve familiarity. Repeat this approach in future testing sessions.
  • Instructions: Explain purpose, procedure, warm-up, number of attempts/trials, scoring method, failed-attempt criteria, and ways to maximize performance.
Administration RuleNumber / Detail
Nonmaximal attemptsRest at least 2 minutes between attempts not close to maximum.
Near-max / max attemptsRest at least 3 minutes between attempts.
Different tests in a batteryRest at least 5 minutes between different tests.
Max anaerobic capacity recoveryGlycolytic system may require 1 hour or more to return to optimal level.
Fatiguing anaerobic and aerobic testsPerform on separate days when possible; if same day, perform last.
Pre-test familiarization1-3 days before actual test at lower exertion level.
Sprint timingUse consistent timing method; anticipate/start delays can create systematic stopwatch error.

11. Logical Sequence of Testing

Core Principle
Testing order should minimize the effect of one test on the next. Do not let fatigue from a glycolytic goblin-test ruin everything that follows.
OrderTest TypeExamples / Notes
1Nonfatiguing testsHeight, weight, body composition, flexibility, vertical jump.
2Agility testsT-test, pro-agility/5-10-5.
3Max power and strength tests1RM/3RM tests, power clean, maximal strength tests.
4Sprint tests10-40 yard/meter sprints.
5Local muscular endurance testsPush-up test, YMCA bench press, prone pull-up.
6Fatiguing anaerobic capacity tests300-yard shuttle, Wingate.
7Aerobic capacity tests1.5-mile run, 12-minute run, Yo-Yo.

12. Test-Specific Notes That Keep Showing Up

Test / MethodHigh-Yield Notes
1.5-mile runMust always be performed last. Validity affected by age, motivation, running familiarity, and aerobic specificity.
12-minute runAerobic capacity estimate; also last because it is fatiguing.
T-test / pro-agilityAgility/change of direction. Should occur before strength and sprint tests.
40-yard dash / sprint testStraight-line speed. Record best of two trials to nearest 0.1 second in many test-bank questions.
Vertical jumpNonfatiguing power test; can occur early. Can also be a sport-specific volleyball/basketball test.
Margaria-KalamenPower test involving stair sprinting.
WingateAnaerobic capacity; highly fatiguing; perform late or separate day.
300-yard shuttleAnaerobic capacity; commonly chosen for court/field intermittent sports.
1RM testsUseful for trained athletes; rest at least 3 minutes between near-max attempts.
3RM / 10RM estimatesCan approximate 1RM when maximal testing is inappropriate. In one quiz-bank item, 10RM was the safer approximation choice.
BESS / SEBT / Y-BalanceBalance and stability tests.
DXAGold standard/comparison test for body composition in many examples.
SkinfoldsPractical body composition field method but may have lower agreement with DXA.
BMILimited for athletes because it does not separate fat from fat-free mass or show fat distribution.

13. Missed-Question Trapboard

These are the common exam traps from my practice questions. Study them twice, then make them do burpees.

TrapCorrect Answer / RuleWhy It Matters
Near-max 1RM attempt rest3 minutes minimum, not 5 minutes.The course protocol gives 3+ minutes for attempts close to maximum.
Measurement vs evaluation vs assessmentMeasurement collects data; evaluation analyzes data to make decisions.Assessment is broader; the test-bank answer for program decisions is evaluation.
Expert judgment of content coverageContent validity, not criterion validity.Criterion-referenced validity is association with another measure.
Hyponatremia clueDilute urine and bloated skin.Heat illness is a major hot-environment risk, but dilute/bloated points to hyponatremia.
Reliability before validityValidity requires consistency.If results are not repeatable, the test cannot validly measure the construct.
Glycolytic recovery after maximal anaerobic capacity1 hour or more, not 3-5 minutes.Do fatiguing anaerobic capacity tests late or on separate days.
Body composition toolsFor body comp, goniometer is not needed.Stadiometer, scale, skinfold calipers, tape/BIA/DXA relate to body comp; goniometer measures ROM.
Max effort attempt rest3 minutes, not 4 minutes.Again: 3+ minutes for near-max/max attempts.
Predictive vs construct validityPredictive = future performance; construct = measures intended construct.Tryout/team-selection questions often point to predictive validity.
Concurrent vs content validityConcurrent validity is a type of criterion validity.Content is expert coverage judgment.
Discriminant vs convergentDiscriminant = low relationship with different constructs; convergent = high relationship with related/gold-standard construct.Percentile ranks across different constructs often test discriminant logic.
Agility vs sprint testsT-test/pro-agility/arrowhead are agility; 12-meter/20-yard sprint tests are speed.Do not label a straight sprint as agility.
1.5-mile run validity factorAge matters in norms/interpretation.Question-bank trap: weight was tempting but age was keyed.
Straight-line sprint trialsBest of two trials.Question-bank trap: you chose three.
Young/low-training athlete strength testAvoid high-risk maximal barbell tests; 1RM dumbbell bench was least recommended in the question-bank item.Technique and safety drive test selection.

14. Quick Compare: Validity Goblin Map

Question Stem Says…Answer
Looks like it measures what it claimsFace validity
Experts judge the test covers all relevant areasContent validity
Measures the intended constructConstruct validity
High correlation with gold standard / related constructConvergent validity
Low correlation with unrelated/different constructDiscriminant validity
Associated with accepted test of same abilityConcurrent validity
Associated with future sport performancePredictive validity

15. Quick Compare: Reliability Goblin Map

Question Stem Says…Answer
Same athlete, same test, different days, similar scoresTest-retest reliability
Different raters/testers disagreePoor interrater reliability
Same tester scores inconsistentlyIntrarater variability
Athlete performs inconsistently across trialsIntrasubject variability
Equipment, biology, and measurement noiseTypical error of measurement
Protocol training for test administrators improves…Interrater reliability

16. Practice Questions and Explanations to Emphasize

Question PatternBest AnswerReason
New sprint test highly correlates with gold-standard timing system.Convergent validityStrong relationship with gold standard measuring same construct.
Underwater weighing aligns with DXA; handheld BIA does not.Concurrent validityAccepted test comparison of same ability/body composition.
Order: pro-agility, 1RM squat, 40-yard dash.Pro-agility -> 1RM squat -> 40-yard dashAgility before max strength; sprint after strength in this simplified set.
Best volleyball test: 40-yard dash, 300-yard shuttle, or 12-minute run.300-yard shuttleMore specific to intermittent anaerobic demands than 12-minute aerobic run.
Scores vary widely under same conditions.Poor reliabilityReliability = consistency.
Vertical jump Monday and Wednesday produce similar scores.Test-retest reliabilityRepeated administrations, same group/athlete.
1RM squat testing across days needs what most?Standardized warm-up and consistent encouragementControls variables directly tied to validity and reliability.

17. One-Page Memory Sheet

BucketMemorize This
PurposeAssess ability and progress; guide goals and program decisions.
MeasurementCollect data.
EvaluationAnalyze data and decide.
ValidityRight thing.
ReliabilitySame thing, repeatedly.
ConstructMeasures intended construct.
FaceLooks valid.
ContentExpert coverage judgment.
ConcurrentAccepted same-ability test, same time.
PredictiveFuture performance.
ConvergentGold standard / related construct, high correlation.
DiscriminantDifferent construct, low correlation.
Rest attempts2+ min not near max; 3+ min near max.
Rest between tests5+ min.
Glycolytic recovery1 hour or more after maximal anaerobic capacity.
Testing orderNonfatiguing, agility, max power/strength, sprint, local muscular endurance, anaerobic capacity, aerobic capacity.
SafetyChest pressure, severe heat signs, confusion, collapse, or unusual symptoms = stop and refer/activate emergency plan.

Leave a comment