Skip to main content
Study NotesLET Secondary · Assessment of LearningReal content

LET Secondary Assessment of LearningPrinciples of Assessment and the Table of SpecificationsStudy Notes

Complete study notes for Principles of Assessment and the Table of Specifications, written for LET Secondary aspirants. Unlike generic notes, these focus on what Professional Regulation Commission (PRC) actually tests in the LET Secondary Assessment of Learning section: high-yield concepts, common question types, and the worked examples that match recent exam patterns.

Exam context

Professional Regulation Commission (PRC) runs the Licensure Examination for Professional Teachers — Secondary on Bi-annual. Its Assessment of Learning section sits under a "Core" weighting, and Principles of Assessment and the Table of Specifications is the 1st chapter in the 5-chapter LET Secondary Assessment of Learning rotation. The LET Secondary passing mark is Weighted average of 75% with no grade below 50%, and the most recent 2026 paper drew about a meaningful share of questions from Assessment of Learning.

Principles of Assessment and the Table of Specifications - Study Notes

Assessment is a cornerstone of effective teaching and learning in the K-6 classroom. As a future elementary teacher in the Philippine education system, you must understand the fundamental principles that guide how we measure, assess, and evaluate student learning. This chapter equips you with the vocabulary, frameworks, and practical tools needed to design fair, valid, and reliable assessments aligned with the Department of Education's K-12 Basic Education Curriculum (BEC) and DepEd Order No. 8, s. 2015 on the Policy Guidelines on Classroom Assessment for the K-12 Basic Education Program. The Licensure Examination for Teachers (LET) tests your mastery of these principles with precision—especially the distinctions between measurement, assessment, and evaluation; the purposes and types of assessment; and the construction of a Table of Specifications as a test blueprint. This foundation ensures you can create assessments that honor your ethical responsibilities under RA 7836 (Code of Ethics for Professional Teachers), which requires teachers to maintain the highest standards of professional competence and to protect the rights of learners.

Summary

Mastering the principles of assessment and the Table of Specifications is foundational to success on the LET and to ethical, effective teaching in Philippine classrooms. **Measurement**, **assessment**, and **evaluation** are three distinct processes that together form a framework for understanding learning. **Assessment purposes**—FOR learning (formative), OF learning (summative), and AS learning (metacognitive)—determine when assessments happen and who benefits. **Types of assessment**—placement, diagnostic, formative, and summative—serve specific functions at different stages of instruction. **Norm-referenced** assessments rank learners relative to peers, while **criterion-referenced** assessments certify mastery against a standard; the DepEd's K-12 Curriculum is fundamentally criterion-referenced. The **Table of Specifications** is your most powerful tool: a blueprint that maps content × cognitive level × item count, ensuring the test is balanced, valid, and fair. It is built using the formula: items = (topic hours / total hours) × total items, and it secures **content validity**, the primary form for teacher-made tests. **Validity**—whether a test measures what it should—comes in five types: content, concurrent, predictive, construct, and face (weakest). **Reliability**—whether a test measures consistently—is established via test-retest, parallel forms, split-half, KR-20, and Cronbach's alpha. The critical relationship: **reliability is necessary but not sufficient for validity; a valid test must be reliable, but a reliable test may not be valid** (tight clustering off-target, like a biased scale). The dartboard analogy captures this perfectly. Under **RA 7836 (Code of Ethics for Professional Teachers)**, educators must maintain the highest standards of professional competence, which includes using assessments that are both valid and reliable. When you build a classroom test using a TOS, write clear items at appropriate cognitive levels, use consistent scoring rubrics, calculate reliability coefficients, and provide constructive feedback, you honor both the science of assessment and the ethical responsibility to students and the teaching profession. The LET will test your mastery of these distinctions, definitions, and applications with precision—especially the vocabulary traps (measurement vs. assessment vs. evaluation, placement vs. diagnostic, norm- vs. criterion-referenced, face vs. construct validity, reliable but not valid). Study the examples closely, practice distinguishing concepts, and internalize the TOS formula and reliability methods. These principles are not just abstract theory; they shape every assessment decision you make as a teacher.

Sections

In Philippine teacher training and on the LET, these three terms are constantly confused, yet each has a precise meaning. Understanding the distinction is essential because many exam items test whether you can apply the right term in context. **Measurement** is the process of **assigning a number or score to a performance or trait**. It answers the question: *How much?* When a student takes a classroom test and earns 42 points out of 50, that is a measurement. The measurement itself is neutral; it simply quantifies what was observed. Measurement is the narrowest of the three processes—it focuses solely on generating a numerical value. In DepEd classrooms, when you record a raw score in your grade book, you are measuring. **Assessment** is the **broader, ongoing collection and interpretation of evidence about learning from multiple sources**. It answers: *What is the evidence of learning?* Assessment goes far beyond a single test score. It includes classroom observations, student interviews, portfolio reviews, project work, quizzes, performance tasks, self-reflections, and peer feedback. Assessment is a holistic, continuous process designed to build a comprehensive picture of what a learner knows and can do. The DepEd's Learning Competency Checks (LCC) and other formative tools are part of classroom assessment. Assessment recognizes that learning is complex and cannot be captured by a single measure. For example, an elementary teacher assesses a Grade 3 student's reading proficiency not just by a reading comprehension test (measurement) but also by observing his fluency during shared reading, examining his written responses to literature, and noting his confidence in reading aloud to peers. **Evaluation** is the **judgment of worth or value made against a standard or criterion**. It answers: *How good? Does this meet the target?* Evaluation interprets the evidence gathered during assessment and makes a decision. When you look at a student's assessment data and decide that a raw score of 42/50 converts to a grade of "Proficient" or "Needs Improvement," you are evaluating. Evaluation is inherently value-laden because it compares performance to a standard—whether that standard is the DepEd's Mastery Benchmarks, a school's grading scale (below average, average, above average), or a criterion set in a lesson plan. Grading, ranking, and certification are all forms of evaluation. Under RA 7836, teachers must ensure that evaluation is fair, objective, and based on clear criteria—not on favoritism or subjective impression. **The Flow**: *Measure → Assess → Evaluate* A practical sequence: You give a test (measurement: assign scores). You collect scores, observations, and work samples (assessment: gather evidence). You then compare the evidence to your learning objectives and decide whether the student has mastered the competency (evaluation: make a judgment). **Classroom Example**: Mrs. Santos, a Grade 4 teacher, is assessing her students' mastery of the competency "Solves problems involving addition and subtraction of fractions with like denominators." She gives a 10-item quiz where students solve word problems (measurement: she records each student's score). She also observes their work in a cooperative learning activity and collects a performance task where pairs solve a real-world problem about sharing pizza slices (assessment: multiple sources of evidence). At the end of the lesson, she reviews all the data—the quiz scores, her observation notes, and the performance task—against the DepEd's Mastery Benchmarks. She decides that 15 of her 40 students have met the competency, 20 are approaching mastery, and 5 need reteaching (evaluation: value judgment against a standard).

Heading

1. Measurement, Assessment, and Evaluation: Three Distinct Processes

Examples

  • A student scores 18/20 on a spelling test = Measurement (raw score assigned)
  • A teacher collects the test score, observation notes, and a writing sample = Assessment (multiple evidence sources)
  • The teacher decides the student is 'Proficient' in spelling = Evaluation (value judgment against a criterion)
  • DepEd's Learning Competency Checks (LCC) yield raw scores (measurement), which are reviewed alongside formative observations (assessment) to determine if the student is ready to progress (evaluation)
  • In a Grade 1 math class, the teacher observes a child counting using concrete objects during a lesson (observation), gives a worksheet with 5 addition problems (measurement), and also watches the child working at a learning station (assessment). The teacher then evaluates: 'Child is ready for abstract number concepts' or 'Needs more concrete practice.'

Key Points

  • Measurement = assigning a number (quantification)
  • Assessment = gathering evidence from multiple sources (holistic process)
  • Evaluation = making a value judgment or decision against a standard
  • The three are sequential: measure, then assess, then evaluate
  • Measurement is objective; assessment is comprehensive; evaluation is interpretive and judgment-based
  • Under RA 7836, teachers must conduct evaluations with integrity and fairness

Modern assessment frameworks—and the LET—emphasize that assessment serves different purposes at different times. These purposes determine *when* the assessment happens, *who uses the results*, and *how the results inform action*. The framework organizes assessment into three linked purposes: FOR learning, OF learning, and AS learning. **Assessment FOR Learning** is **formative assessment**. It occurs **during instruction** to provide feedback that helps teachers adjust their teaching and helps learners adjust their learning while there is still time to intervene. The primary user of the information is the *teacher* (and secondarily the learner). The question being answered is: *Are students progressing? What do they understand, and what do they still need?* Formative assessments are typically **ungraded or low-stakes**. Examples include quick oral checks during a lesson ("Who can show me with your fingers how many halves make a whole?"), exit slips, think-pair-share activities, informal observation, and classroom quizzes. The DepEd's emphasis on **daily formative assessment** as part of the Normal Learning Activities (NLA) reflects this purpose. When Mrs. Santos pauses her Grade 3 math lesson and asks students to hold up number cards showing their answer to an addition problem, she is using formative assessment FOR learning—to gauge class understanding in real time and decide whether to move forward or re-teach. **Assessment OF Learning** is **summative assessment**. It occurs **at the end of a unit, quarter, or school year** to certify and report what students have achieved. It answers: *Have the intended learning outcomes been met?* The primary user is external stakeholders—parents, school leaders, the next teacher, or the system (via standardized tests). Summative assessments are **graded and high-stakes**. They include unit tests, quarterly assessments, the National Achievement Test (NAT), and school division tests. They provide the final evidence of mastery and inform grades. Under DepEd policy, summative assessments should account for performance in periodic tests (40%), written work and products (40%), and performance tasks and skills (20%), as per DepEd's grading guidelines. **Assessment AS Learning** develops the learner's **metacognition** and **self-regulation**. It happens when students **monitor their own progress, reflect on their learning, and take ownership of improvement**. The primary user is the *learner themselves*. Examples include student self-assessment checklists ("I can read this with fluency: yes/no/sometimes"), learning logs or journals where students reflect ("What was hard today? What helped me?"), peer assessment rubrics, and student conferences where children articulate their understanding. This purpose is aligned with the DepEd's 21st-century skills framework, which emphasizes learner agency and metacognition. When a Grade 5 student uses a self-assessment rubric to evaluate his own essay against the criteria of organization, voice, and mechanics, he is engaging in assessment AS learning—he becomes aware of his strengths and gaps and is motivated to revise. **How They Work Together**: A **balanced assessment system** uses all three. A teacher gives formative checks during a lesson (FOR), a quiz at the end of a week (OF), and asks students to rate their own confidence in the skill on a Likert scale (AS). This three-fold approach aligns with DepEd policy and reflects best practice. The LET often presents a scenario—a teacher does X after a lesson—and asks which purpose this represents. Your job is to identify whether the goal is *immediate feedback to adjust teaching* (FOR), *certification of mastery* (OF), or *student self-awareness* (AS). **Classroom Example**: In a Grade 2 reading lesson on phonics, Mrs. De la Cruz uses assessment FOR learning when she stops mid-lesson and has students blend sounds aloud (she listens and notes who is struggling). She uses assessment OF learning when she gives a phonics test at the end of the week and records the scores in her grade book. She uses assessment AS learning when she gives each child a smiley-face self-check: "I can blend sounds: 😊 (yes) / 😐 (so-so) / ☹ (not yet)." Together, these three purposes paint a full picture of the child's phonics progress.

Heading

2. Purposes of Assessment: FOR, OF, and AS Learning

Examples

  • A teacher gives a 2-minute quiz on multiplication facts mid-lesson and immediately reviews errors with the class = FOR learning (formative feedback)
  • A teacher administers a unit test on fractions at the end of the week and records the grades = OF learning (summative, graded)
  • A student checks off items on a rubric: 'I understood the lesson: yes/mostly/no' = AS learning (self-assessment, metacognition)
  • DepEd's learning competency checks during regular instruction = FOR learning; the summative test at the end of the quarter = OF learning
  • A Grade 4 student writes in a learning journal: 'Today I found it hard to write my paragraph because I didn't know how to start. Tomorrow I will ask my teacher for help.' = AS learning (reflection and self-regulation)
  • A teacher observes a child's grip on a pencil during a writing lesson and provides corrective feedback on the spot = FOR learning; a formal handwriting assessment submitted at the end of the month for grading = OF learning

Key Points

  • Assessment FOR Learning = Formative, during instruction, feedback to adjust teaching and learning, low-stakes
  • Assessment OF Learning = Summative, end of unit/term, certification of mastery, graded, high-stakes
  • Assessment AS Learning = Student self-assessment and metacognition, learner is the primary user, builds agency
  • FOR is about feedback and adjustment; OF is about grading and reporting; AS is about student awareness and motivation
  • DepEd policy and the LET expect teachers to use all three purposes in a balanced assessment system
  • Assessment FOR learning is aligned with DepEd's emphasis on formative assessment during daily NLA

Beyond the purposes just outlined, assessments are also classified by **when they are given** and **what specific function they serve**. This classification yields four types: placement, diagnostic, formative, and summative. Many LET items hinge on distinguishing placement from diagnostic, so pay close attention. **Placement Assessment** occurs **before instruction begins** to determine **entry-level readiness and where to position the learner**. It answers: *What does this student already know? Where should instruction start?* Placement assessments are not meant to diagnose *why* a gap exists; they simply determine the appropriate starting point. Examples include a pre-test given on the first day of a unit, screening tools used at the start of the school year, or a kindergarten readiness screening. The information is used to **group students by readiness level** (homogeneous grouping for small-group instruction) or to **adjust the starting point of instruction**. A Grade 3 teacher might give a phonics pre-test to see who is ready to move to more advanced word families and who needs to consolidate blending; this is placement. Placement is forward-looking: it guides the road ahead. **Diagnostic Assessment** occurs **before instruction (pretest) or during instruction** to **pinpoint specific strengths, weaknesses, learning difficulties, and their underlying causes**. It answers: *What exactly is the learner struggling with, and why?* Diagnostic assessments are deeper and more individualized than placement. They probe the nature of errors and misconceptions. When a Grade 2 student consistently reverses the digits 6 and 9, a diagnostic assessment might include one-on-one observation, interviews, and specialized tasks to explore whether the issue is visual discrimination, motor control, spatial awareness, or conceptual confusion. Diagnostic assessments are used to **design targeted interventions** and to identify learners who may need referral for special education evaluation. Under DepEd policy, teachers use diagnostic tools (such as the Individual Pupil Monitoring System [IPMS] or Learning Difficulty Checklist) to flag students requiring additional support or specialist consultation. Diagnostic assessment is detective work: it seeks to uncover the root of the problem. **Formative Assessment** occurs **during instruction** to **monitor progress and provide ongoing feedback**. (This is the same as Assessment FOR Learning.) It is typically **ungraded or carries minimal weight in final grades**. Examples include observation checklists, short quizzes, exit slips, and think-pair-share responses. Its function is immediate feedback that allows the teacher to adjust instruction and students to refine their understanding in real time. **Summative Assessment** occurs **at the end of a unit, quarter, or school year** to **judge and certify overall achievement against learning objectives**. (This is the same as Assessment OF Learning.) It is **graded and weighted heavily** in final grades. Examples include unit tests, quarterly exams, NAT, and final projects that are evaluated against a rubric and entered into the grade book. **Key Distinction: Placement vs. Diagnostic** The LET often asks: "A teacher gives a pretest to find out where to start instruction. Is this placement or diagnostic?" The answer is **placement**. Placement simply maps readiness; it does not probe for the cause of gaps. If the question had said "to understand *why* the student is behind," the answer would be **diagnostic**. Placement asks, "Are you ready?"; diagnostic asks, "What's the problem, and what caused it?" **Classroom Example**: At the start of a Grade 4 unit on division, Ms. Reyes gives a 5-question pre-test on multiplication (her students' prior knowledge). Students who score 4–5 are ready for division with larger numbers; those who score 2–3 need a brief review of multiplication; those who score 0–1 need intensive remediation. This is **placement**—it positions each student at the right starting point. However, when Ms. Reyes notices that one child, Juan, scored 1, she sits down with him one-on-one and works through a few multiplication problems aloud. She observes that Juan knows his facts but rushes and makes careless errors (reads a 3 as a 5). She also notices he has no multiplication chart to reference. Ms. Reyes now understands the *cause* of Juan's gaps—not lack of knowledge, but careless reading and lack of a support tool. This deeper investigation is **diagnostic**. She will provide Juan with a multiplication chart and a speed-check strategy to slow him down. The placement information got Juan to the right level; the diagnostic information revealed how to help him succeed at that level.

Heading

3. Types of Assessment: Placement, Diagnostic, Formative, and Summative

Examples

  • A Grade 1 teacher gives a letter-sound identification test on the first day = Placement (determines starting point for phonics instruction)
  • A Grade 4 teacher notices a student solving word problems incorrectly and does a one-on-one interview to discover the student misunderstands the concept of 'remainder' = Diagnostic (probes the root cause)
  • A Grade 5 teacher gives a brief spelling quiz every Friday to see who mastered the week's words = Formative (ongoing, low-stakes feedback)
  • A Grade 3 teacher administers a comprehensive reading test at the end of the quarter and records the grade = Summative (certified achievement, graded)
  • DepEd's Brigada Eskwela screening activities for children entering Grade 1 = Placement (readiness check)
  • A teacher's use of the IPMS to identify students at risk of not learning to read = Diagnostic (targeted intervention design)

Key Points

  • Placement assessment = before instruction, determines readiness and starting point, forward-looking
  • Diagnostic assessment = before or during instruction, pinpoints difficulties and root causes, targeted intervention
  • Formative assessment = during instruction, provides ongoing feedback, ungraded or low-stakes
  • Summative assessment = end of unit/term, judges overall achievement, graded and high-stakes
  • Placement answers 'Are you ready?'; Diagnostic answers 'What's wrong and why?'
  • DepEd expects teachers to use diagnostic tools (IPMS, Learning Difficulty Checklists) to identify learners needing intervention
  • The LET frequently tests the placement-vs-diagnostic distinction

This distinction concerns the **standard against which performance is compared and interpreted**. It is crucial for the LET because it affects how scores are reported and what they mean. **Norm-Referenced Assessment (NRA)** interprets performance **relative to other test-takers in a comparison group (the norm group)**. It answers: *How does this person rank compared to others?* A norm-referenced test deliberately includes items of varying difficulty to **spread out scores** and create a distribution so that some students score high, most score in the middle, and some score low. The goal is to **discriminate and rank learners**, not to certify mastery of specific skills. Results are reported as **percentile ranks, stanines, grade equivalents, or relative rankings** (e.g., "You scored higher than 85% of your class"; "Your child ranks in the 72nd percentile nationally"). Examples of norm-referenced tests include the National College Entrance Examination (NCEE, which feeds the UPCAT), the Licensure Examination for Teachers (LET) itself, and classroom tests designed to rank students from highest to lowest achiever. The Philippine National Achievement Test (NAT), while nominally criterion-referenced, has often been interpreted in norm-referenced ways (comparing schools and regions to each other). A norm-referenced interpretation does **not** guarantee that a student has mastered specific competencies—only that they performed better or worse than peers. **Criterion-Referenced Assessment (CRA)** interprets performance **against a fixed, predetermined standard or criterion (the mastery level)**. It answers: *Has this person met the target?* A criterion-referenced test is designed to match **specific learning objectives**, and items are selected because they measure those objectives, not because they have particular difficulty levels. The goal is to **certify whether the learner has mastered the content**, regardless of how others performed. Results are reported as **mastery/non-mastery, proficient/not yet proficient, or a percentage of competencies met** (e.g., "Learner can divide fractions by fractions: Yes/No"; "Student achieved 85% mastery of the learning competencies for Grade 3 Mathematics"). Examples include classroom quizzes aligned to DepEd learning competencies, driving tests (pass/fail against the traffic code), and teacher-made unit assessments with a clear mastery benchmark. DepEd's K-12 Curriculum Framework and all learning competency checks are fundamentally **criterion-referenced**: the question is whether a student has demonstrated the competency, not whether they scored better than classmates. **The Relationship to the LET**: While the LET uses a cut score (75% weighted average), it is functionally **criterion-referenced**. You pass if you meet the criterion (75%); you do not pass based on ranking. However, the LET also reports your standing relative to other examinees (e.g., "You placed in the top 10% of examinees"), which is a norm-referenced layer. The **primary interpretation is criterion-referenced**: Does the examinee demonstrate minimal teaching competence? The ranking is secondary information. **Key Comparison Table**: | Aspect | Norm-Referenced | Criterion-Referenced | | --- | --- | --- | | Compared against | Other test-takers (group norm) | Fixed standard or criterion | | Item selection | Varied difficulty to spread scores | Matched to objectives, difficulty varies with objective | | Purpose | Rank and discriminate learners | Certify mastery of competencies | | Reporting | Percentile, stanine, grade equivalent, ranking | Percentage mastery, proficient/not proficient, competency achieved/not achieved | | Example | NCEE, entrance exams, rank-based tests | DepEd learning competency assessments, mastery quizzes, skills certifications | | Implication | High score = better than peers; low score = worse than peers | Score shows whether the criterion is met; peers' scores are irrelevant | **Why This Matters in the Philippine Classroom**: The DepEd's shift to competency-based education means that **most classroom assessments should be criterion-referenced**. A Grade 3 teacher assessing "Writes simple sentences with correct subject-verb agreement" against a rubric is using criterion-referenced logic: either the sentence shows agreement, or it doesn't. The fact that another student did worse or better is not the point. However, teachers also use norm-referenced thinking when they look at a class's test scores, see a wide distribution, and decide to rank students or group them by achievement level. Both approaches have a place, but the DepEd's primary framework is criterion-referenced. **Classroom Example**: Mr. Tanaka gives his Grade 5 class a multiplication test. On a **norm-referenced interpretation**, he looks at the scores (ranging from 42/50 to 98/50), sees that Juan scored 88/50, and says, "Juan did very well—he's in the top 10% of the class." The focus is on Juan's rank. On a **criterion-referenced interpretation**, Mr. Tanaka decides in advance that 85/50 or higher shows mastery of the learning competency "Multiplies numbers up to 4 digits by numbers up to 2 digits." He reviews Juan's score and says, "Juan has achieved the competency." The focus is on whether the criterion (85/50) is met. DepEd policy expects Mr. Tanaka to use the criterion-referenced interpretation: either Juan has mastered multiplication (yes) or he hasn't (no), based on a pre-set benchmark, not on how his classmates did.

Heading

4. Norm-Referenced versus Criterion-Referenced Assessment

Examples

  • A student receives a percentile rank of 92 on a national standardized test = Norm-referenced (ranks against all test-takers)
  • A student is told 'You have not yet achieved the learning competency in reading comprehension' based on a mastery benchmark = Criterion-referenced (compared to standard)
  • A Grade 4 teacher ranks her class by math test scores from highest to lowest = Norm-referenced interpretation
  • A Grade 4 teacher marks each student's essay 'Proficient' or 'Developing' against a detailed rubric = Criterion-referenced
  • The LET's passing criterion: 75% weighted average (parts not lower than 50%) = Criterion-referenced; LET's announcement of top 10 passers = Norm-referenced layer
  • DepEd's report on 'Percentage of Grade 6 students nationwide achieving mastery in Mathematics' = Criterion-referenced (mastery is the standard)
  • A Grade 3 teacher notes that a child scored in the 65th percentile on the NAT = Norm-referenced interpretation (ranked against the nation)
  • The same teacher notes 'Child achieved 7 out of 10 learning competencies in Grade 3 Mathematics' = Criterion-referenced (competency count)

Key Points

  • Norm-referenced compares performance to a group (ranking, percentiles)
  • Criterion-referenced compares performance to a fixed standard (mastery yes/no)
  • NRT aims to discriminate and rank; CRT aims to certify and report mastery
  • NRT reports relative standing; CRT reports mastery or competency achievement
  • DepEd's K-12 Curriculum is criterion-referenced; most classroom assessments should use CRT logic
  • The LET is primarily criterion-referenced (75% cut score) but also reports norm-referenced rankings
  • Examples: UPCAT = NRT; DepEd learning competency checks = CRT

The **Table of Specifications (TOS)** is a **two-way chart that serves as a test blueprint**. It maps **content topics** (rows) against **cognitive levels or types of learning objectives** (columns) and specifies the **number of test items** for each cell. The TOS is perhaps the single most important tool for ensuring a test is **content-valid** and that it reflects the emphasis and complexity of the instruction it measures. Every teacher who creates a classroom test should construct a TOS, and many LET items test your understanding of how to build one. **Why a TOS Matters** 1. **Ensures Content Validity**: The TOS guarantees that the test samples the subject matter adequately and proportionally. No topic is over-represented (inflating its apparent importance) or under-represented (making results misleading). This is the most direct way to document **content validity**. 2. **Aligns with Learning Objectives**: By mapping objectives to cognitive levels, the TOS ensures items target the intended depth of thinking. If your goal is to develop application skills, the TOS shows that items emphasize the Apply level, not just Remember and Understand. 3. **Balances Difficulty**: The TOS helps distribute items across difficulty levels (easy, moderate, hard) in a planned way. The DepEd's Enhanced Table of Specifications recommends roughly **30% easy (Remember/Understand), 50% moderate (Apply/Analyze), and 20% difficult (Evaluate/Create)** items. 4. **Communicates Test Design to Stakeholders**: A documented TOS shows students, parents, and administrators the fairness and thoughtfulness behind a test. It answers the question: "Why are these particular items on the test?" 5. **Guides Item Writing**: The TOS tells item writers which content and cognitive level each item must address, preventing random item selection. **How to Construct a TOS: Step-by-Step** **Step 1: List the Content Topics** Identify all the major topics covered in the unit or term. For a Grade 3 Mathematics unit on multiplication, topics might be: - Skip counting by 2s, 5s, and 10s - Multiplication as repeated addition - Multiplication tables (2x, 5x, 10x) - Solving word problems with multiplication - Mental computation strategies **Step 2: State Learning Objectives and Cognitive Levels** For each topic, identify the learning objective and classify it by Bloom's Revised Taxonomy: Remember, Understand, Apply, Analyze, Evaluate, or Create. (Elementary tests rarely reach Evaluate or Create; most focus on Remember through Apply.) Example: - Topic: Multiplication tables — Objective: "Recall the 2x, 5x, and 10x multiplication facts" — Cognitive level: **Remember** - Topic: Solving word problems — Objective: "Apply multiplication to solve real-world problems" — Cognitive level: **Apply** **Step 3: Determine Total Number of Items** Decide how many items the test will contain. For a Grade 3 math unit quiz, 20 items might be appropriate. For a quarterly assessment, 40–50 items. **Step 4: Allocate Items by Content Topic** Distribute items across topics in proportion to the **instructional time or emphasis** each received. Use the formula: **Items per topic = (Hours spent on topic / Total hours) × Total items** Example: If a unit is taught over 15 hours: - Skip counting (2 hours) → (2/15) × 20 = 2.67 ≈ 3 items - Multiplication as repeated addition (3 hours) → (3/15) × 20 = 4 items - Multiplication tables (6 hours) → (6/15) × 20 = 8 items - Word problems (3 hours) → (3/15) × 20 = 4 items - Mental strategies (1 hour) → (1/15) × 20 = 1.33 ≈ 1 item - **Total: 20 items** Notice that the allocation reflects the time invested: multiplication tables got the most time (6 hours) and the most items (8). This ensures the test mirrors the instruction. **Step 5: Distribute Items by Cognitive Level** Within each topic's items, divide them across cognitive levels. For elementary assessments, follow the **30/50/20 distribution**: - 30% at Remember/Understand (easy): recall facts, identify concepts - 50% at Apply/Analyze (moderate): solve problems, explain reasoning - 20% at Evaluate/Create (difficult): justify choices, design solutions For the multiplication tables (8 items), this might be: - 2–3 items at Remember ("What is 5 × 6?") — easy - 4–5 items at Apply ("Maria has 3 bags of apples. Each bag has 10 apples. How many apples does she have?") — moderate - 1–2 items at Analyze/Evaluate ("Show two different ways to solve 4 × 7 and explain which way is faster for you") — difficult **The Table Itself** A TOS is typically formatted as a matrix: | Content Topic | Remember (%) | Understand (%) | Apply (%) | Analyze (%) | Total Items | | --- | --- | --- | --- | --- | --- | | Skip counting | 2 | 1 | — | — | 3 | | Repeated addition | 1 | 1 | 2 | — | 4 | | Multiplication tables | 3 | — | 4 | 1 | 8 | | Word problems | — | — | 3 | 1 | 4 | | Mental strategies | — | — | 1 | — | 1 | | **Total** | **6 (30%)** | **2** | **10 (50%)** | **2 (20%)** | **20** | Each cell shows the number of items at that cognitive level for that topic. The bottom row totals show the overall cognitive distribution: 30% Remember/Understand, 50% Apply, 20% Analyze. **Why This Matters for the LET**: The LET itself is constructed using an Enhanced Table of Specifications with four subscales (Math, Science, English, Social Studies) and a difficulty distribution. Your job as a test builder is to use the same logic. An LET item might ask: "A teacher is constructing a test on photosynthesis. She spent 40% of class time on the structure of the chloroplast, 35% on the light reactions, 20% on the Calvin cycle, and 5% on applications. If the test has 50 items, how many items should address the light reactions?" The answer is (35/100) × 50 = 17.5 ≈ 18 items. This is the **allocation formula**. **Worked Example: A Grade 5 English Test on Expository Writing** Suppose a Grade 5 teacher is designing a 30-item test on expository writing. She taught: - Paragraph structure (3 hours) - Topic sentences (2 hours) - Supporting details (4 hours) - Transitions (2 hours) - Editing and revising (4 hours) - Total: 15 hours Step 4—Allocate items by topic: - Paragraph structure: (3/15) × 30 = 6 items - Topic sentences: (2/15) × 30 = 4 items - Supporting details: (4/15) × 30 = 8 items - Transitions: (2/15) × 30 = 4 items - Editing and revising: (4/15) × 30 = 8 items Step 5—Distribute by cognitive level (30/50/20): - Paragraph structure (6 items): 2 Remember, 3 Apply, 1 Analyze - Topic sentences (4 items): 1 Remember, 2 Apply, 1 Analyze - Supporting details (8 items): 2 Remember, 5 Apply, 1 Analyze - Transitions (4 items): 1 Remember, 2 Apply, 1 Analyze - Editing and revising (8 items): 3 Remember, 4 Apply, 1 Analyze Total: 9 items at Remember (30%), 16 items at Apply (53%), 5 items at Analyze (17%) ≈ 30/50/20. Now the teacher writes items aligned to each cell. For example: - A Remember item on paragraph structure: "A paragraph has how many main parts?" - An Apply item on paragraph structure: "Read this paragraph. Identify the topic sentence, supporting details, and concluding sentence." - An Analyze item on supporting details: "Which set of supporting details best supports the topic sentence 'Recycling helps the environment'?" The TOS ensures the test is balanced, fair, and valid. **Alignment with DepEd**: The DepEd's Enhanced Table of Specifications for the LET follows this same logic. It specifies how many items appear in each subject, at each difficulty level, ensuring that no area is overemphasized and that the test reflects the breadth and depth of teacher preparation required by the K-12 Curriculum Framework.

Heading

5. The Table of Specifications (TOS): Blueprint for Valid Testing

Examples

  • A teacher plans a 40-item science test. She spent 10 hours on the water cycle, 8 hours on weather, and 12 hours on ecosystems. For the water cycle: (10/30) × 40 = 13 items; weather: (8/30) × 40 = 11 items; ecosystems: (12/30) × 40 = 16 items.
  • A Grade 4 math test has 20 items: 6 on addition facts (easy), 8 on solving word problems with addition (moderate), and 2 on multi-step reasoning (difficult) = 30/40/10 distribution (roughly 30/50/20).
  • A teacher writes a TOS for a Grade 3 science unit on animals. Rows: habitats, food chains, adaptation. Columns: Remember, Understand, Apply. The TOS shows 2 items on Remember/adaptation, 3 items on Apply/food chains, etc., totaling 20 items, each mapped to a learning competency.
  • An LET question: "A teacher allocates items on a test using this formula: (time on topic / total time) × total items. She spent 5 days on fractions and 10 days on decimals, and wants a 30-item test. How many items on decimals? Answer: (10/15) × 30 = 20 items."
  • A kindergarten teacher builds a TOS for a letter-recognition assessment: Topics—uppercase letters, lowercase letters, sound-symbol connection. Cognitive levels—Remember (identify letter shown), Understand (match upper to lower). The TOS has 20 items, with 12 on Remember and 8 on Understand, aligned to DepEd's learning competencies for Grade K.

Key Points

  • Table of Specifications = test blueprint mapping content × cognitive level × number of items
  • TOS ensures content validity and proportional representation of topics
  • TOS guides item writing and ensures alignment with learning objectives
  • Allocation formula: Items per topic = (Hours for topic / Total hours) × Total items
  • Follow 30/50/20 rule: 30% easy (Remember/Understand), 50% moderate (Apply/Analyze), 20% difficult (Evaluate/Create)
  • A documented TOS demonstrates fairness and thoughtfulness in test design
  • The LET uses an Enhanced TOS; teachers should apply the same logic to classroom tests
  • DepEd policy expects tests to be aligned to learning competencies via a TOS

**Validity** is the **most important quality of any assessment**. It is the degree to which a test measures what it is intended to measure and the appropriateness and usefulness of the interpretations and decisions made from the test scores. A test can be reliable (consistent) without being valid (measuring the right thing), but a valid test must be reliable. Validity is not an all-or-nothing property; it exists on a continuum. Different types of validity serve different purposes and are established through different methods. **Content Validity** **Content validity** is the degree to which a test adequately and representatively samples the content domain it claims to measure. It answers: *Does the test cover the breadth and depth of the subject matter as it was taught?* For a Grade 4 mathematics test on addition and subtraction, content validity means that the test includes problems of various types (with and without regrouping), various contexts (money, measurement, word problems), and items at the appropriate cognitive levels. Content validity is established through **expert judgment and a Table of Specifications**. A panel of teachers, curriculum specialists, or subject-matter experts reviews the test items and the TOS and judges whether the items fairly represent the content taught. There is no numerical formula for content validity; it is qualitative judgment grounded in the TOS. For the LET, you are expected to know that a TOS is the primary tool for documenting and securing content validity. If an LET question says, "A teacher constructs a test without a TOS to guide item selection. What aspect of test quality is at risk?" the answer is **content validity**. Content validity is also the most relevant type for teacher-made classroom tests. **Criterion-Related Validity** **Criterion-related validity** is the degree to which test scores relate to an external criterion—a measure that is presumed to be a true reflection of the construct. It is established through **correlation: comparing test scores to the criterion and seeing how closely they align**. Criterion-related validity splits into two subtypes depending on *when* the criterion is measured. **Concurrent Validity** is established when the criterion is measured **at the same time as the test**. It answers: *Do scores on this test align with a current, parallel measure of the same construct?* For example, a teacher develops a new, quick 5-question fluency quiz. To check concurrent validity, she compares students' scores on the quick quiz with their scores on a longer, validated reading fluency assessment given the same day. If correlations are high (say, 0.85 or higher), the quick quiz has good concurrent validity—it gives the same information as the longer test. Concurrent validity is useful for developing efficient screening tools. **Predictive Validity** is established when the criterion is measured **in the future**. It answers: *Do current test scores predict future performance on an important outcome?* For instance, do Grade 3 Math scores predict Grade 5 Math scores? Do LET scores predict teaching effectiveness or student learning gains in the first year of teaching? The LET's predictive validity would ideally be demonstrated by showing that higher LET scorers become more effective teachers (though this is complex to measure and track). Predictive validity is critical for entrance exams like the UPCAT, which must show that entrance exam scores predict college success (graduation rates, GPA). If UPCAT scores did not correlate with college performance, the UPCAT's predictive validity would be questioned. Criterion-related validity is quantified using **correlation coefficients** (typically Pearson's r, ranging from −1 to +1). A coefficient of 0.70 or higher is often considered acceptable; 0.80+ is strong. Criterion-related validity is more relevant for standardized tests and entrance exams than for classroom formative assessments. **Construct Validity** **Construct validity** is the degree to which a test measures the theoretical **construct** it claims to measure—that is, the underlying trait or concept. It answers: *Does this test truly measure what we think it measures (e.g., intelligence, anxiety, reading comprehension)?* Constructs are abstract. "Reading comprehension," for instance, is not directly observable; we infer it from test responses. Construct validity is established through multiple lines of evidence, including: - **Convergent evidence**: Scores on this test correlate highly with other measures of the same construct (e.g., a new reading comprehension test correlates with a validated reading comprehension measure). - **Divergent evidence**: Scores on this test do *not* correlate highly with measures of unrelated constructs (e.g., a reading comprehension test should not correlate strongly with a spelling test, which measures a different skill). - **Factor analysis**: Statistical technique to identify the underlying structure and confirm that the test measures a single, unitary construct. - **Experimental evidence**: Tests to see whether scores change as expected when the construct is manipulated. Construct validity is the most complex type and is usually relevant for standardized ability and personality tests, not classroom quizzes. **Face Validity** **Face validity** is the degree to which a test *appears* to measure what it claims to measure—that is, the test's superficial appearance. It is **the weakest form of validity** because appearance is not the same as actual measurement. A test can have high face validity (look reasonable on the surface) but low actual validity (fail to measure the intended construct). For example, a "creativity test" that asks students to list as many uses for a brick as possible has face validity (it looks like it measures creativity), but research suggests that fluency (quantity of responses) is not the same as true creativity (originality, flexibility, elaboration). Educators and test-takers expect a test to have face validity because it builds confidence and motivation, but face validity should never be confused with actual validity. The LET often includes a question like: "A test appears to measure critical thinking but actually measures only recall. This test has low ___ validity." The answer is **criterion-related** or **construct** validity, not face validity—because face validity means it *looks* like it measures something, which this test does. The issue is that the actual measurement (what it *really* measures) does not match the claim. **Summary: Types of Validity and Their Purposes** | Type | Definition | How it's checked | Relevance | | --- | --- | --- | --- | | **Content** | Adequate, representative sampling of the domain | Expert judgment, TOS | Teacher-made tests (primary) | | **Concurrent** | Scores align with current criterion | Correlation with concurrent measure | Screening tools, efficiency studies | | **Predictive** | Scores predict future criterion | Correlation with future performance | Entrance exams, placement tests | | **Construct** | Test measures the theoretical trait | Convergent/divergent evidence, factor analysis | Ability/personality tests | | **Face** | Superficial appearance of validity | Inspection | Secondary; does not confirm actual validity | **Why Validity Matters in the Philippine Classroom** Under RA 7836 (Code of Ethics for Professional Teachers), educators must "maintain the highest standards of professional competence and excellence." Using valid assessments is part of that responsibility. A teacher who creates a test without a TOS and thus lacks content validity is not meeting ethical standards—she may be grading students unfairly. DepEd's emphasis on learning competency-based assessment also hinges on validity: the learning competencies define the domain, and a valid test samples that domain proportionally. A Grade 4 teacher assessing "Reads with fluency: Grade 4 level" must include varied texts (fiction, nonfiction, poetry) and contexts (independent reading, oral reading, paired reading) to have content validity. If she tests only poetry, the test lacks validity for measuring the broader competency. **Classroom Example**: Mrs. Santos develops a test to measure her Grade 3 students' "mastery of two-digit addition." She writes 15 items—all basic facts like 23 + 14 = ?, with no regrouping required. The test has low content validity because it does not represent the full domain taught (addition WITH regrouping, word problems, different number contexts). A more valid test would include problems with and without regrouping, word problems, and mental math checks. By using a TOS that maps the content taught and allocating items proportionally, Mrs. Santos would secure content validity. Later, when Mrs. Santos wants to validate a new, quicker addition quiz, she correlates its scores with students' scores on a longer, established addition test given the same week. If the correlation is high (0.80+), the quiz has **concurrent validity** and can be used as a time-efficient screener.

Heading

6. Validity: Does the Test Measure What It Should?

Examples

  • A Grade 3 math test has high content validity if its TOS shows items covering all topics taught (addition, subtraction, word problems, mental math) in proportion to instructional time.
  • A new reading fluency measure correlates 0.82 with an established fluency assessment given the same day = High concurrent validity (works as well as the established measure).
  • UPCAT scores correlate 0.65 with first-year college GPA = Predictive validity evidence (entrance scores predict college performance).
  • A test of 'critical thinking' correlates highly with another 'critical thinking' test (convergent) but does not correlate with a spelling test (divergent) = Construct validity evidence.
  • A 'creativity test' asks students to list uses for a pencil. It has face validity (looks like it measures creativity) but low construct validity if research shows quantity of responses (what it actually measures) is not the same as true creativity.
  • A teacher assesses 'listening comprehension' using only written multiple-choice items. Students who struggle with reading comprehension may score low even if their listening comprehension is strong. The test lacks construct validity; it measures reading comprehension more than listening.
  • DepEd's learning competency assessments are designed with high content validity: each item is mapped to a specific competency, and the TOS ensures all competencies are sampled proportionally.

Key Points

  • Validity = the degree to which a test measures what it claims to measure; it is the most important quality of a test
  • Content validity = adequate, representative sampling of the domain; established via TOS and expert judgment; primary type for teacher-made tests
  • Criterion-related validity = scores relate to an external criterion
  • - Concurrent validity = criterion measured at the same time (current alignment)
  • - Predictive validity = criterion measured in the future (forecast of performance)
  • Construct validity = test measures the underlying abstract trait; established via convergent/divergent evidence and factor analysis
  • Face validity = superficial appearance of validity; weakest form; does not confirm actual validity
  • A test can be reliable but not valid; a valid test must be reliable
  • RA 7836 requires teachers to maintain high standards; valid assessment is part of that ethical responsibility

**Reliability** is the **consistency, stability, or dependability of test scores**. A reliable test gives similar results under similar conditions. Reliability is about precision and repeatability. If you weigh yourself on a scale three times in a row and get three different weights (158 lbs, 162 lbs, 160 lbs), the scale is unreliable—it is not precise. In education, if a student takes a test and scores 42/50 one day but scores 28/50 when taking an equivalent test the next week (assuming no learning or forgetting occurred), the test is unreliable—the scores are inconsistent. Reliability is necessary for validity but is not sufficient; a test can be reliable (consistent) without being valid (measuring the right thing). **Methods of Establishing Reliability** Reliability can be established through several approaches, each measuring a different type of consistency. **Test-Retest Reliability** **Test-retest reliability** is established by administering the same test to the same group on two separate occasions and **correlating the scores from the two administrations**. It measures **stability of scores over time**. The formula used is typically **Pearson's correlation coefficient (r)**. If test-retest reliability is high (e.g., r = 0.85), scores are stable: people who scored high on the first administration also scored high on the second, and those who scored low also scored low. This method is most appropriate for stable constructs (knowledge, skills) and less so for constructs that naturally fluctuate (mood, fatigue). **Limitations**: - **Practice effect**: Taking the test the first time may improve performance on the second administration simply because learners are familiar with the format or remember answers. - **Memory effect**: Conversely, memory of the first test might inflate correlation artificially. - **Time elapsed**: If too much time passes, learning or forgetting may occur, artificially lowering correlation. - **Demand on participants**: Asking people to take a test twice is burdensome and may reduce cooperation. **Parallel (Equivalent) Forms Reliability** **Parallel forms reliability** is established by creating two **equivalent versions of a test** (Form A and Form B), administering both to the same group, and **correlating the two sets of scores**. It measures **equivalence of forms**. The two forms must measure the same constructs, have the same difficulty level, the same number of items, and the same time limit. If parallel forms reliability is high (e.g., r = 0.83), both forms are equally valid measures of the construct. **Advantages**: - Avoids practice and memory effects because items are different. - Allows the use of two different forms for pre- and post-testing without confounding effects. **Disadvantages**: - Requires development of two equally valid, equivalent tests—time-consuming and difficult to accomplish. - Still burdens participants with two test administrations. **Split-Half Reliability** **Split-half reliability** is established by **splitting a single test into two halves** (typically odd-numbered items vs. even-numbered items, or first half vs. second half), calculating the score for each half, and **correlating the two half-scores**. It measures **internal consistency**: the degree to which different parts of the test correlate, reflecting whether items measure the same construct. However, splitting a test into halves creates two shorter tests, and **shorter tests are less reliable**. To correct for this, the **Spearman-Brown prophecy formula** is applied to "correct" or "inflate" the correlation to what it would be if the full test were two full tests: **r_full = (2 × r_half) / (1 + r_half)** For example, if the correlation between odd and even items is 0.75, the Spearman-Brown corrected reliability is: (2 × 0.75) / (1 + 0.75) = 1.5 / 1.75 = 0.857. **Advantages**: - Requires only one administration—quick and economical. - No practice or memory effects. **Disadvantages**: - Only reflects internal consistency at one moment in time; does not assess stability over time or equivalence across forms. - Assumes the two halves are truly equivalent (may not hold if difficulty varies across the test). **Internal Consistency (Cronbach's Alpha and KR-20/KR-21)** **Internal consistency** measures how much test items **intercorrelate**—how much items measuring the same construct agree with each other. High internal consistency suggests that all items are measuring the same underlying trait and are homogeneous. Two common statistics are used: 1. **Cronbach's Alpha (α)**: Used for tests with items on a **scale** (e.g., Likert-type items: strongly disagree, disagree, agree, strongly agree). Alpha ranges from 0 to 1; values of 0.70+ are generally considered acceptable, and 0.80+ are strong. 2. **Kuder-Richardson Formula 20 (KR-20)**: Used for tests with **dichotomous items** (right/wrong, yes/no). Like alpha, KR-20 ranges from 0 to 1. 3. **Kuder-Richardson Formula 21 (KR-21)**: A simplified version of KR-20 that does not require item-by-item analysis; instead, it uses only the mean and variance of total scores. **Advantages**: - Requires only one administration. - Directly reflects item homogeneity. - Modern statistical software calculates these automatically. **Disadvantages**: - High internal consistency may indicate that items are too similar, reducing construct diversity. For example, a critical thinking test with all items on identifying assumptions might have high alpha but does not measure all aspects of critical thinking. **A Summary Table of Reliability Methods** | Method | Procedure | Type of consistency | Advantages | Disadvantages | | --- | --- | --- | --- | --- | | **Test-retest** | Same test, same group, two times | Stability over time | Direct measure of stability | Practice/memory effects, time burden | | **Parallel forms** | Two equivalent tests, same group | Equivalence of forms | No practice/memory effects | Requires two equivalent tests, time burden | | **Split-half + Spearman-Brown** | One test split in half, correlation corrected | Internal consistency | One administration, quick | Time-limited; assumes equivalent halves | | **KR-20 / Alpha** | Statistical analysis of item intercorrelations | Internal consistency | One administration, modern, direct | High alpha may indicate over-similarity of items | **Interpreting Reliability Coefficients** Reliability coefficients range from 0 (no consistency) to 1 (perfect consistency). General guidelines: - **0.90+**: Excellent reliability; suitable for high-stakes decisions (e.g., entrance exams, licensing exams like the LET) - **0.80–0.89**: Good reliability; acceptable for most classroom tests - **0.70–0.79**: Fair reliability; useful for initial screening or low-stakes measures - **Below 0.70**: Poor reliability; generally not recommended for decision-making The LET aims for reliability coefficients of 0.90+ to ensure that passing decisions are based on stable, consistent measurement. **Why Reliability Matters in the Classroom** Consider a Grade 4 math test. If the test has low reliability, a student might score 35/50 on Monday but 48/50 on Wednesday (assuming no learning occurred), simply because the test items were ambiguous or the scoring was inconsistent. This unreliability could lead to unfair grading decisions. Under DepEd policy, teachers must ensure their assessments are reliable. A reliable test: - Produces consistent scores for the same learner under similar conditions. - Allows teachers to trust their grading decisions. - Ensures fairness: learners are evaluated consistently, not subject to random variation. - Provides meaningful feedback: trends in scores reflect actual learning changes, not measurement error. **Worked Example**: A Grade 5 teacher creates a 20-item reading comprehension test. To check internal consistency, she uses KR-20. The calculation (done by software) yields KR-20 = 0.82. This indicates that the 20 items are highly consistent in measuring reading comprehension; the test is reliable. When she gives the test to her students, she can trust that their scores reflect their actual comprehension abilities, not random variation or confusion from poorly written items. **Reliability and the LET**: The LET is extensively normed and validated. Its reliability coefficients are published and typically exceed 0.90, ensuring that passing scores reliably reflect teaching competence. As a test-taker, you should know that the LET is a reliable instrument; your score is a stable, dependable measure of your readiness to teach. **Sources of Unreliability (Error)** Understanding sources of unreliability helps teachers reduce it. Common sources include: - **Ambiguous items**: Poorly written questions that students interpret differently. - **Inconsistent scoring**: Teacher grades subjectively or applies rubrics inconsistently. - **Guessing**: On multiple-choice tests, correct guesses inflate scores unreliably. - **Test anxiety and fatigue**: A learner's mental state during testing affects performance inconsistently. - **Environmental factors**: Noise, interruptions, or discomfort lower scores unpredictably. - **Item heterogeneity**: Items measure different constructs, so total score is muddled. By writing clear items, using detailed rubrics, reducing guessing (e.g., via constructed-response items), managing test conditions, and ensuring items focus on a single construct, teachers increase reliability.

Heading

7. Reliability: Does the Test Measure Consistently?

Examples

  • A teacher gives a 25-item reading test on Monday and the same test on Friday to the same Grade 4 class. Correlating the two sets of scores yields r = 0.87 = High test-retest reliability (scores are stable over the week).
  • A teacher develops two equivalent 20-item math tests (Form A and Form B), gives both to her Grade 3 class on the same day, and correlates the scores: r = 0.84 = Good parallel forms reliability.
  • A teacher splits a 40-item social studies test into odd items (20) and even items (20), calculates the score on each half, correlates them: r_half = 0.70. Spearman-Brown correction: (2 × 0.70) / (1 + 0.70) = 0.824 = Good split-half reliability.
  • A teacher uses statistical software to calculate Cronbach's alpha for a 15-item attitude survey: α = 0.89 = Excellent internal consistency; the items are highly intercorrelated and measure a single attitude construct.
  • A Grade 6 teacher's history test has alpha = 0.68 (poor reliability). She reviews items and finds they measure multiple, unrelated topics (geography, dates, cultural context). She rewrites to focus narrowly on 'Timeline of Philippine independence' and recalculates alpha = 0.81 (good). Internal consistency improved by making items more homogeneous.
  • The LET reports reliability of 0.94 (excellent). This assures test-takers that their scores are stable and trustworthy; passing reflects genuine teaching competence, not random measurement error.

Key Points

  • Reliability = consistency and stability of test scores; a reliable test gives similar results under similar conditions
  • Test-retest reliability = stability over time; one group, test twice, correlate scores
  • Parallel forms reliability = equivalence of forms; two equivalent versions, same group, correlate
  • Split-half reliability = internal consistency; split one test in half, correlate halves, apply Spearman-Brown correction
  • KR-20 / Cronbach's Alpha = statistical internal consistency; calculate from one administration
  • Reliability coefficient ranges 0–1; 0.90+ = excellent, 0.80–0.89 = good, 0.70–0.79 = fair, below 0.70 = poor
  • Reliability is necessary for validity but not sufficient; a reliable test may not be valid
  • Sources of unreliability include ambiguous items, inconsistent scoring, guessing, test anxiety, and environmental factors
  • DepEd expects teachers to ensure assessment reliability; the LET has reliability coefficients of 0.90+

One of the most tested concepts on the LET is the relationship between validity and reliability. Many educators confuse or conflate the two, but they are distinct properties, and their relationship has important implications. **The Key Principle**: **A test can be reliable without being valid, but a valid test must be reliable.** In other words: - **Reliability is necessary but not sufficient for validity.** - **Validity implies reliability.** This is not intuitive for many test-takers, so let's unpack it carefully with the famous **dartboard analogy**. **The Dartboard Analogy** Imagine a target with a bullseye at the center and concentric rings around it. You throw darts (test scores) at the target (the construct you want to measure). 1. **Arrows clustered tightly in the CENTER of the bullseye** = **Valid AND Reliable** - Tight clustering = Reliable (consistent, stable scores) - Centered on bullseye = Valid (measuring the right thing) - Ideal test: scores cluster around the true measure; you can trust them 2. **Arrows clustered tightly OFF-CENTER from the bullseye** = **Reliable but NOT Valid** - Tight clustering = Reliable (consistent, predictable) - Off-center = Not valid (not measuring the right thing; systematic bias) - Example: A scale that always reads 2 pounds too high. If you step on it three times, it gives 125 lbs, 125 lbs, 125 lbs (reliable). But your true weight is 123 lbs, so the scale is biased and not valid. The measurements are consistent but wrong. - In education: A test of mathematical reasoning that always measures only arithmetic facts (reliable—consistent measurement of facts) but not reasoning (invalid—not measuring the target construct). 3. **Arrows scattered all over the target** = **Neither Reliable NOR Valid** - Scattered = Unreliable (inconsistent, all over the place) - No clustering near bullseye = Invalid (not measuring the target; too much random error) - Example: A bathroom scale that reads 120 lbs one day, 135 lbs the next, 118 lbs the next. Unreliable and unreliable, thus not useful. - In education: A poorly written test with ambiguous items, inconsistent scoring, and high guessing. Scores are unpredictable and do not reflect true ability. **Why Reliability Is Necessary But Not Sufficient for Validity** Think of it this way: If you are aiming at a target and your throws are scattered all over (unreliable), you cannot possibly hit the bullseye consistently (invalid). **You must first be precise (reliable) before you can be accurate (valid).** Reliability is the precondition. However, being precise does not guarantee accuracy. You might be precisely off-target (clustered off-center), as in the biased scale example. Thus: - **Unreliable → Cannot be valid** (no precision, so no accuracy possible) - **Reliable but off-target → Possible** (precise but wrong) - **Valid → Must be reliable** (if you are measuring the right thing consistently, you are being precise) **Classroom Example**: Mrs. Garcia develops a test of "critical thinking in science." All 30 items ask students to recall facts ("What is the chemical formula for water?"). The test is **reliable**: questions are clear, scoring is objective (right/wrong), and students' scores are consistent when given the test twice. KR-20 = 0.88 (high internal consistency). However, the test is **not valid** for measuring critical thinking because critical thinking requires analysis, evaluation, and synthesis—not just fact recall. The test is reliable (precise, consistent) but off-target (not measuring the intended construct). To improve validity, Mrs. Garcia must rewrite items to include "Why...?" and "How...?" questions that require analysis and reasoning. **Why This Matters for the LET** The LET tests this distinction directly with items like: "A teacher's test has high internal consistency (KR-20 = 0.87) but students score low on authentic performance tasks assessing the same skill. The test most likely has low ___ validity." The answer is **construct validity** or **criterion-related validity** (depending on phrasing)—because the test is reliable (high KR-20) but not valid (scores don't align with authentic performance). Another question: "A test is valid. What can you assume about its reliability?" The answer is **the test is also reliable**—because validity implies reliability. **Ensuring Both Reliability and Validity** To create a test that is both reliable and valid: 1. **Build a Table of Specifications** (ensures content validity). 2. **Write clear, unambiguous items** (improves reliability and construct validity). 3. **Use consistent scoring procedures** (improves reliability). 4. **Pilot test and calculate reliability coefficients** (confirms reliability). 5. **Correlate scores with criterion measures** (establishes criterion-related or construct validity). 6. **Review by experts** (confirms content validity). Under RA 7836, teachers are ethically bound to use assessments that are both valid (measuring the right learning outcomes) and reliable (producing consistent, trustworthy scores). An unreliable test is unfair to students; an invalid test may be grading skills that were not taught or that are not relevant to the learning objectives.

Heading

8. The Relationship Between Validity and Reliability: A Critical Distinction

Examples

  • A test of reading comprehension has KR-20 = 0.85 (reliable) but does not correlate with students' performance in literature class (low criterion-related validity) = Reliable but not valid.
  • A critical thinking test is designed with a TOS, reviewed by experts for content validity, has KR-20 = 0.89, and scores correlate 0.82 with students' performance on authentic reasoning tasks = Valid and reliable.
  • A handwriting assessment rubric is subjectively applied by a teacher (different standards applied to different papers) = Low reliability (inconsistent scoring) → Cannot be valid for handwriting.
  • A spelling test is highly consistent (KR-20 = 0.91) but all items test only vowel-consonant patterns, ignoring silent letters, suffixes, and common exceptions taught in class = Reliable but low content validity.
  • An LET question: 'A classroom test has poor internal consistency but high correlation with a validated standardized test. This test most likely has low ___ and high ___ validity.' Answer: low **internal consistency/reliability**, high **criterion-related validity**. The test is not reliable but is valid because it correlates with a criterion.

Key Points

  • Reliability is necessary but not sufficient for validity
  • A valid test is necessarily reliable
  • A test can be reliable but not valid (precise but off-target)
  • A test cannot be valid but unreliable (accuracy requires precision)
  • Dartboard analogy: tight-and-centered = valid+reliable; tight-but-off-center = reliable only; scattered = neither
  • Use all three methods to ensure both: TOS (validity), clear items (reliability and construct validity), consistent scoring (reliability), expert review (validity)
  • RA 7836 requires teachers to use assessments that are both valid and reliable

Now that you understand the principles of measurement, assessment, evaluation, purposes, types, validity, and reliability, let's synthesize them into a practical, step-by-step process for building a classroom test that honors both principles and DepEd policy. This is essential preparation for the LET because many items ask you to identify flaws in test design or to recommend improvements. **Step 1: Define Learning Objectives and Cognitive Levels** Before writing a single test item, state clearly what students should know and be able to do by the end of the unit. Reference the DepEd's K-12 Curriculum learning competencies. For a Grade 4 English unit, an objective might be: "Students will identify the main idea and supporting details in an expository text (Apply level)." This objective guides everything that follows. **Step 2: Build a Table of Specifications** Create a two-way table: - **Rows**: Content topics covered in the unit - **Columns**: Cognitive levels (Remember, Understand, Apply, Analyze) - **Cells**: Number of items for each topic × cognitive level Allocate items using the formula: (instructional time on topic / total time) × total items. Follow the 30/50/20 distribution (30% easy, 50% moderate, 20% difficult). This TOS is your blueprint and is the foundation of **content validity**. **Step 3: Write Items Aligned to the TOS** For each cell in the TOS, write items at that cognitive level on that topic. Use these guidelines: - **Remember items**: Require simple recall or identification. Example: "What is the main idea of this paragraph?" - **Understand items**: Require explaining or summarizing. Example: "Explain how the supporting details connect to the main idea." - **Apply items**: Require using knowledge in a new context. Example: "Read this new paragraph and identify the main idea." - **Analyze items**: Require comparing, contrasting, or examining reasoning. Example: "How do these two paragraphs' main ideas relate?" Ensure items are **clear and unambiguous** to reduce random error and improve reliability. **Step 4: Determine Scoring and Grading Procedures** Decide how items will be scored: - **Objective items** (multiple-choice, matching): Clear right/wrong answers; score 1 point or 0 points. - **Constructed-response items** (short answer, essay): Use a detailed **rubric** with clear criteria and point values. For example, an essay might be scored 0–4 on "Organization," 0–4 on "Content," 0–4 on "Mechanics," yielding a total of 0–12 points. Using a detailed rubric **improves reliability** by reducing scorer bias and inconsistency. Share the rubric with students before the test so they know expectations. **Step 5: Pilot Test (if possible) and Calculate Reliability** If feasible, give the test to a small group before using it for grading. Calculate **internal consistency** using KR-20 (for objective items) or Cronbach's alpha (for subjective items). Aim for reliability ≥ 0.80 for classroom assessments. If reliability is low (< 0.70): - Review items for clarity and remove ambiguous ones. - Ensure all items measure the same construct (not multiple unrelated topics). - Review scoring consistency. **Step 6: Administer Under Controlled Conditions** Reduce sources of unreliability: - Minimize distractions and interruptions. - Allow adequate time (no time pressure that artificially lowers scores). - Ensure all students have the same materials and instructions. - Monitor to prevent cheating. **Step 7: Score Consistently and Review Results** Apply the scoring rubric uniformly to all students. After scoring, analyze results: - Are there items that most students missed? These might be poorly written or too difficult. - Are there items that everyone got right? These might be too easy and could be replaced with harder items next time. - Do scores follow a reasonable distribution (some low, some high, cluster in the middle for a classroom test)? This review informs future test improvement. **Step 8: Provide Constructive Feedback** **Assessment FOR Learning** happens here. Return the test with specific feedback: "You understood the main idea (✓) but struggled to connect it to supporting details. Let's practice that skill together." Feedback should be actionable and tied to learning objectives, not just a grade. **Alignment with DepEd and RA 7836** This process aligns with DepEd's expectations: - **Learning Competency Alignment**: The TOS maps to DepEd's published competencies, ensuring that your assessment reflects the official curriculum. - **Balanced Assessment**: By using multiple cognitive levels, you assess depth of understanding, not just recall—aligned with 21st-century skills. - **Formative-Summative Balance**: The test can serve as a summative assessment (graded) and also generate formative feedback (improvement guidance). - **Ethical Practice**: Under RA 7836, you are maintaining professional standards by using valid, reliable assessments. This is fair to students and honors their right to clear, consistent evaluation. **Common Pitfalls to Avoid** 1. **No TOS**: Writing items without a blueprint. Result: Unbalanced test, possible low content validity. 2. **All recall items**: Overloading the test with Remember-level items. Result: Does not assess higher-order thinking despite teaching it. 3. **Ambiguous items**: Poorly worded questions that students interpret differently. Result: Unreliable scoring and frustration. 4. **Inconsistent rubrics**: Grading essays with no criteria or applying different standards to different students. Result: Low reliability, perceived unfairness. 5. **No feedback**: Returning only a grade. Result: Missed opportunity for formative assessment (FOR learning). 6. **Guessing-prone design**: All multiple-choice, no constructed response. Result: Scores inflated by lucky guesses, unreliable estimate of true ability. 7. **No reliability check**: Never calculating internal consistency. Result: Unknown reliability; may be grading unfairly without realizing it. **Classroom Example**: Mr. Ramos is preparing a Grade 5 Science unit test on "The Water Cycle." He identifies his learning objectives (based on DepEd competencies), builds a TOS mapping content (evaporation, condensation, precipitation, collection) × cognitive levels (Remember through Analyze), with a total of 30 items distributed as 9 at Remember (30%), 15 at Apply (50%), and 6 at Analyze (20%). He writes items aligned to each cell. For constructed-response items, he develops a 4-point rubric (incomplete, developing, proficient, advanced). He gives the test to his class, calculates KR-20 for objective items (0.82—good) and alpha for rubric-scored items (0.86—good). He identifies that one item was confusing (half the class missed it; the wording was unclear) and will revise it next time. He returns tests with written feedback: "Great understanding of the water cycle's stages! Next, work on explaining WHY evaporation happens, not just WHAT it is." The test is valid (aligned to learning objectives, TOS-based, expert-reviewed), reliable (calculated coefficients > 0.80), and instructionally useful (formative feedback provided).

Heading

9. Practical Synthesis: Building a Valid and Reliable Classroom Test

Examples

  • A Grade 3 reading test: TOS rows = phonics, fluency, comprehension; columns = Remember, Apply; teacher allocates 10 items to phonics Remember, 5 to phonics Apply, 8 to comprehension Apply, etc. Each item is written to match its TOS cell.
  • A Grade 4 math test uses a rubric for word problems: 0 pts = no work shown; 1 pt = work shown but incorrect answer; 2 pts = correct answer with some work; 3 pts = correct answer with complete, clear work. Teacher applies this rubric uniformly to all students, then calculates alpha = 0.85 (reliable rubric).
  • A Grade 6 social studies test includes multiple-choice items (scored 1 or 0) and an essay (scored 0–5 via rubric). Teacher calculates KR-20 for MC items (0.79), alpha for essay (0.83), and overall internal consistency (0.81). All three indicate acceptable reliability.
  • A Grade 5 teacher gives students the TOS and rubric before the test. Students know 40% of the test is about applying fractions to real-world problems (Apply level) and 20% is comparing and analyzing fraction strategies (Analyze level). This transparency supports learning.
  • A teacher reviews her Grade 4 test results. She notices 85% of students missed item 15, while 92% got item 8 correct. She re-reads item 15 and sees it's poorly worded. She removes it and replaces it with a clearer version next semester. Item analysis improves test quality.

Key Points

  • Build a TOS before writing items; allocate based on instructional time; follow 30/50/20 difficulty distribution
  • Write clear items aligned to each TOS cell; use detailed rubrics for constructed responses
  • Calculate reliability (KR-20, alpha); aim for ≥ 0.80 for classroom assessments
  • Administer under controlled conditions; score consistently; provide constructive feedback
  • Review results to identify confusing items and plan improvements
  • Avoid pitfalls: no TOS, all recall, ambiguous items, inconsistent rubrics, no feedback, guessing-prone design, no reliability check
  • Alignment with DepEd K-12 Curriculum and RA 7836 ethical standards is essential
Loading diagram…
Loading diagram…
Loading diagram…
Loading diagram…
Loading diagram…

Ready to practise for the LET Secondary 2026?

Super Tutor's AI review plan adapts to your weak areas and builds a weekly practice schedule around your target LET Secondary exam date.