Skip to content · ⁨본문 바로가기⁩

Collecting Data · ⁨데이터 수집⁩

AP Statistics · ⁨AP 통계학⁩ · Topic 3 · ⁨주제 3⁩

Video lesson for this topic · ⁨이 주제용 영상 수업⁩ Open the video page · ⁨영상 페이지 열기⁩
7:31

Collecting Data · ⁨데이터 수집⁩

An online poll collects thirty thousand replies. A careful survey asks only one thousand people. Which is closer to the truth? The small one, almost every…

English narration · English + 中文 subtitles burned in · ⁨영어 내레이션 · 영어 + 중국어 자막 burned-in⁩

3.1

Can We Trust the Data We Collected? · ⁨우리가 수집한 데이터를 신뢰할 수 있을까요?⁩

Syllabus
English

Enduring Understanding (VAR-1): Given that variation may be random or not, conclusions are uncertain.

Learning Objective VAR-1.E: Identify questions to be answered about data collection methods. [Skill 1.A]

  • VAR-1.E.1 Methods for data collection that do not rely on chance result in untrustworthy conclusions.
한국어

지속적 이해 (VAR-1): 변동성이 우연일 수도 있고 그렇지 않을 수도 있으므로, 결론은 불확실할 수 있다.

학습 목표 VAR-1.E: 데이터 수집 방법에 대해 답해야 할 질문을 식별하기. [기술 1.A]

  • VAR-1.E.1 확률에 의존하지 않는 데이터 수집 방법은 신뢰할 수 없는 결론을 초래한다.

Source: College Board AP Course and Exam Description · ⁨출처: College Board AP Course and Exam Description⁩

English

A conclusion is only as good as the data behind it. How data are collected decides what you may conclude – whether you can generalize to a population 总体, and whether you can claim cause and effect. Poorly collected data can be worse than none.

한국어

결론은 그에 기반한 데이터의 품질만큼이나 좋습니다. 데이터 수집 방법이 결론의 범위(보편 집단(population)에 일반화 가능한지, 인과관계를 주장할 수 있는지를 결정합니다. 잘못 수집된 데이터는 아예 없는 것보다 나쁠 수 있습니다.

Vocabulary · ⁨어휘⁩ Train · ⁨연습하기⁩
English 한국어
population/ˌpɒpjʊˈleɪʃn/ 인구
3.2

Observational Studies and Experiments · ⁨관찰 연구와 실험⁩

Syllabus
English

Enduring Understanding (DAT-2): The way we collect data influences what we can and cannot say about a population.

Learning Objective DAT-2.A: Identify the type of a study. [Skill 1.C]

  • DAT-2.A.1 A population consists of all items or subjects of interest.
  • DAT-2.A.2 A sample selected for study is a subset of the population.
  • DAT-2.A.3 In an observational study, treatments are not imposed. Investigators examine data for a sample of individuals (retrospective) or follow a sample of individuals into the future collecting data (prospective) in order to investigate a topic of interest about the population. A sample survey is a type of observational study that collects data from a sample in an attempt to learn about the population from which the sample was taken.
  • DAT-2.A.4 In an experiment, different conditions (treatments) are assigned to experimental units (participants or subjects).

Learning Objective DAT-2.B: Identify appropriate generalizations and determinations based on observational studies. [Skill 4.A]

  • DAT-2.B.1 It is only appropriate to make generalizations about a population based on samples that are randomly selected or otherwise representative of that population.
  • DAT-2.B.2 A sample is only generalizable to the population from which the sample was selected.
  • DAT-2.B.3 It is not possible to determine causal relationships between variables using data collected in an observational study.
한국어

영구적 이해 (DAT-2): 데이터를 수집하는 방식은 우리가 집단에 대해 말할 수 있고 말할 수 없는 것을 결정한다.

학습 목표 DAT-2.A: 연구의 유형을 식별하기. [기술 1.C]

  • DAT-2.A.1 집단은 관심 있는 모든 항목이나 주체로 구성된다.
  • DAT-2.A.2 연구에 선택된 표본은 집단의 하집합이다.
  • DAT-2.A.3 관찰 연구에서는 처리가 가해지지 않는다. 연구자는 집단에 대한 관심 있는 주제에 대해 조사하기 위해 개인 표본의 데이터를 검토(후향적)하거나 개인 표본을 미래로 추적하여 데이터를 수집(전향적)한다. 표본 조사는 표본이 추출된 집단에 대해 알기 위해 표본으로부터 데이터를 수집하려는 관찰 연구의 일종이다.
  • DAT-2.A.4 실험에서는 서로 다른 조건(처리)이 실험 단위(참여자 또는 피험자)에 배정된다.

학습 목표 DAT-2.B: 관찰 연구를 기반으로 한 적절한 일반화 및 판단을 식별하기. [기술 4.A]

  • DAT-2.B.1 집단에 대한 일반화는 랜덤하게 선택되었거나 해당 집단을 대표하는 표본에 대해서만 적절하다.
  • DAT-2.B.2 표본은 그 표본이 선택된 집단에서만 일반화 가능하다.
  • DAT-2.B.3 관찰研究中에서 수집된 데이터를 사용하여 변수 간의 인과 관계를 결정하는 것은 불가능하다.

Source: College Board AP Course and Exam Description · ⁨출처: College Board AP Course and Exam Description⁩

English
  • In an observational study 观察性研究 you measure individuals without trying to influence them. It can show association, but not causation, because lurking variables may explain the link.
  • In an experiment 实验 you deliberately impose a treatment 处理 and compare responses. A well-designed experiment can establish cause and effect.
한국어
  • **관찰 연구(observational study)**에서는 개인을 측정하되 영향을 주지 않습니다. **연관성(association)**을 보일 수는 있으나, 숨겨진 변수가 연결을 설명할 수 있으므로 인과관계는 알 수 없습니다.
  • **실험(experiment)**에서는 의도적으로 **처치(treatment)**를 가하고 반응을 비교합니다. 잘 설계된 실험은 인과관계를 확립할 수 있습니다.
Explore · ⁨탐색하기⁩

Observational study or experiment? · ⁨관찰 연구인지 실험인가?⁩

In an experiment the researcher imposes a treatment (and can show cause); an observational study only records what already happens (and can show association, not cause). · ⁨실험에서 연구자는 처방을 가하여 인과관계를 증명할 수 있으며, 관찰 연구는 이미 발생하는 현상을 기록할 뿐이며(인과관계가 아닌 상관관계만 보여줍니다).⁩

Vocabulary · ⁨어휘⁩ Train · ⁨연습하기⁩
English 한국어
observational study/ɒbzəˈveɪʃənl ˈstʌdi/ 관찰 연구
experiment/ekˈsperɪmənt/ 실험
treatment/ˈtriːtmənt/ 처치
3.3

Random Sampling · ⁨무작위 표본 추출⁩

Syllabus
English

Enduring Understanding (DAT-2): The way we collect data influences what we can and cannot say about a population.

Learning Objective DAT-2.C: Identify a sampling method, given a description of a study. [Skill 1.C]

  • DAT-2.C.1 When an item from a population can be selected only once, this is called sampling without replacement. When an item from the population can be selected more than once, this is called sampling with replacement.
  • DAT-2.C.2 A simple random sample (SRS) is a sample in which every group of a given size has an equal chance of being chosen. This method is the basis for many types of sampling mechanisms. A few examples of mechanisms used to obtain SRSs include numbering individuals and using a random number generator to select which ones to include in the sample, ignoring repeats, using a table of random numbers, or drawing a card from a deck without replacement.
  • DAT-2.C.3 A stratified random sample involves the division of a population into separate groups, called strata, based on shared attributes or characteristics (homogeneous grouping). Within each stratum a simple random sample is selected, and the selected units are combined to form the sample.
  • DAT-2.C.4 A cluster sample involves the division of a population into smaller groups, called clusters. Ideally, there is heterogeneity within each cluster, and clusters are similar to one another in their composition. A simple random sample of clusters is selected from the population to form the sample of clusters. Data are collected from all observations in the selected clusters.
  • DAT-2.C.5 A systematic random sample is a method in which sample members from a population are selected according to a random starting point and a fixed, periodic interval.
  • DAT-2.C.6 A census selects all items/subjects in a population.

Learning Objective DAT-2.D: Explain why a particular sampling method is or is not appropriate for a given situation. [Skill 1.C]

  • DAT-2.D.1 There are advantages and disadvantages for each sampling method depending upon the question that is to be answered and the population from which the sample will be drawn.
한국어

영구적 이해 (DAT-2): 데이터를 수집하는 방식은 우리가 집단에 대해 말할 수 있고 말할 수 없는 것을 결정한다.

학습 목표 DAT-2.C: 연구 설명을 바탕으로 샘플링 방법을 식별하기. [기술 1.C]

  • DAT-2.C.1 집단의 항목이 한 번만 선택될 수 있을 때는 이를 복원/sample without replacement이라고 한다. 집단의 항목이 여러 번 선택될 수 있을 때는 이를 부복원/sample with replacement이라고 한다.
  • DAT-2.C.2 단순랜덤표본(SRS)은 특정 크기의 모든 그룹이 선택될 동일한 확률을 가지는 표본이다. 이 방법은 다양한 샘플링 메커니즘의 기초가 된다. SRS를 얻기 위해 사용되는 몇몇 메커니즘의 예로는 개인에게 번호를 매겨 랜덤넘버생성기를 사용하여 표본에 포함시킬 것들을 선택(중복 무시), 랜덤넘버표 사용, 또는 부복원으로 카드를 뽑는 등이 있다.
  • DAT-2.C.3 층화 무작위 표본(stratified random sample)은 공통된 속성이나 특징(동질적 그룹화)에 따라 모집단을 개별적인 그룹인 층(strata)으로 나누는 것이다. 각 층 내부에서 단순 무작위 표본이 선택되며, 선택된 단위들이 결합되어 최종 표본을 구성한다.
  • DAT-2.C.4 클러스터 표본(cluster sample)은 모집단을 작은 그룹인 클러스터(clusters)로 나누는 방법이다. 이상적으로는 각 클러스터 내부에 이질성이 존재하며, 클러스터끼리 구성 면에서 유사해야 한다. 클러스터의 단순 무작위 표본을 모집단에서 선택하여 클러스터 표본을 형성한다. 선택된 클러스터 내의 모든 관측치로부터 데이터를 수집한다.
  • DAT-2.C.5 체계적 무작위 표본(systematic random sample)은 무작위 시작점과 고정된 주기적 간격을 사용하여 모집단의 표본 구성원을 선택하는 방법이다.
  • DAT-2.C.6 전수 조사(census)는 모집단에 있는 모든 항목/대상을 선정한다.

학습 목표 DAT-2.D: 특정 situations에 대한 sampling method가 적절한지 아닌지를 설명할 수 있다. [Skill 1.C]

  • DAT-2.D.1Sampling method마다 답하고자 하는 질문과 표본을 추출할 모집단에 따라 장단점이 다르다.

Source: College Board AP Course and Exam Description · ⁨출처: College Board AP Course and Exam Description⁩

English

To learn about a population you take a sample 样本. Random sampling 随机抽样 protects against selection bias 偏差 and lets you generalize (it cannot fix undercoverage, nonresponse, or response bias — see below). Common designs:

  • Simple random sample (SRS) 简单随机样本: every group of the chosen size is equally likely.
  • Stratified 分层: split the population into similar strata, then sample within each.
  • Cluster 整群: split into clusters, randomly choose whole clusters.
  • Systematic 系统: pick every $k$th individual from a random start.

A convenience sample 方便样本 or voluntary response sample is not random and is biased.

Worked example. To survey a school, an administrator lists all students by grade and randomly selects $20$ from each grade. This is a stratified sample – the grades are the strata – which guarantees every grade is represented, unlike an SRS that might by chance draw few from one grade.

한국어

보편 집단에 대해 배우려면 **표본(sample)**을 취합니다. **무작위 표본 추출(random sampling)**은 선택 **편향(bias)**을 방지하고 일반화를 가능하게 합니다(아래에서 설명할 저면_coverage, 비응답, 응답 편향을 해결하지는 못함). 일반적인 설계:

  • 简单随机样本 (SRS): 선택한 크기의 모든 그룹이 동일한 확률로 선택됩니다.
  • 층화(stratified): 집단을 유사한 층(strata)으로 나누어 각 층 내에서 표본을 취합니다.
  • 군집(cluster): 클러스터로 나누어 전체 클러스터를 무작위로 선택합니다.
  • 체계적(systematic): 무작위 시작점에서 매 $k$번째 개인을 선택합니다.
네 가지 무작위 표본 추출 설계: 누가 선택되고 어떻게 하는지
네 가지 무작위 표본 설계: 누가 선정되며, 어떻게

편의 표본 또는 자발적 응답 표본은 무작위가 아니며, 편향되어 있습니다.

해설 예시. 학교를 조사하기 위해 관리자가 학년별students을 나열하고 각 학년에서 $20$명을 무작위로 선택합니다. 이는 층화 표본입니다 – 학년이 층(strata)이며 – 모든 학년이 반드시 포함됨을 보장합니다. 이는 우연히 한 학년부터 소수의 표본이 추출될 수 있는 SRS와 다릅니다.

무작위 결과: 공정한 조건에서 주사위는 각 면이 동등하게 나올 수 있음
무작위 결과: 공정한 조건에서 주사위는 각 면이 동등하게 나올 수 있음
Vocabulary · ⁨어휘⁩ Train · ⁨연습하기⁩
English 한국어
sample/ˈsæmpl/ 표본
Random sampling/ˈrændəm ˈsæmplɪŋ/ 무작위 표본 추출
bias/ˈbaɪəs/ bias
Simple random sample (SRS)/ˈsɪmpl ˈrændəm ˈsæmpl/ 简单随机样本 (SRS) → 단순무작위표본 (SRS)
Stratified/ˈstrætɪfaɪd/ 층화抽样
Cluster/ˈklʌstə/ 군집抽样
Systematic/ˌsɪstəˈmætɪk/ 체계적抽样
convenience sample/kənˈviːnɪəns ˈsæmpl/ 편의 표본
3.4

When Sampling Goes Wrong · ⁨표본 추출 시 오류 발생⁩

Syllabus
English

Enduring Understanding (DAT-2): The way we collect data influences what we can and cannot say about a population.

Learning Objective DAT-2.E: Identify potential sources of bias in sampling methods. [Skill 1.C]

  • DAT-2.E.1 Bias occurs when certain responses are systematically favored over others.
  • DAT-2.E.2 When a sample is comprised entirely of volunteers or people who choose to participate, the sample will typically not be representative of the population (voluntary response bias).
  • DAT-2.E.3 When part of the population has a reduced chance of being included in the sample, the sample will typically not be representative of the population (undercoverage bias).
  • DAT-2.E.4 Individuals chosen for the sample for whom data cannot be obtained (or who refuse to respond) may differ from those for whom data can be obtained (nonresponse bias).
  • DAT-2.E.5 Problems in the data gathering instrument or process result in response bias. Examples include questions that are confusing or leading (question wording bias) and self-reported responses.
  • DAT-2.E.6 Non-random sampling methods (for example, samples chosen by convenience or voluntary response) introduce potential for bias because they do not use chance to select the individuals.
한국어

영구적 이해 (DAT-2): 데이터를 수집하는 방식은 우리가 집단에 대해 말할 수 있고 말할 수 없는 것을 결정한다.

학습 목표 DAT-2.E: Sampling methods에서 발생할 수 있는 bias의 잠재적 원인을 식별할 수 있다. [Skill 1.C]

  • DAT-2.E.1 bias는 특정 응답이 다른 응답보다 체계적으로 선호될 때 발생한다.
  • DAT-2.E.2 표본이 전적으로 자발적으로 참여하거나 참여를 선택한 사람들로만 구성되어 있을 경우, 그 표본은 일반적으로 모집단을 대표하지 못한다(자발 응답 편향, voluntary response bias).
  • DAT-2.E.3 모집단의 일부가 표본에 포함될 확률이 낮아졌을 경우, 그 표본은 일반적으로 모집단을 대표하지 못한다(미포함 편향, undercoverage bias).
  • DAT-2.E.4 데이터 확보가 불가능하거나(또는 응답을 거부한) 표본에 선정한 개인들은 데이터 확보가 가능한 개인들과 다를 수 있다(비응답 편향, nonresponse bias).
  • DAT-2.E.5 데이터 수집 도구나 과정의 문제가 반응 편향(response bias)을 유발한다. 예로는 혼란스럽거나 유도적인 질문(질문 문구 편향, question wording bias), 자기 보고식 응답 등이 포함된다.
  • DAT-2.E.6 비무작위 sampling methods(예: 편의抽样或自愿响应抽样)는 chances를 사용하여 individual들을 선택하지 않기 때문에 bias가 발생할 가능성이 있다.

Source: College Board AP Course and Exam Description · ⁨출처: College Board AP Course and Exam Description⁩

English

Bias makes estimates systematically miss the truth:

  • Undercoverage 覆盖不足: some groups are left out of the sampling frame.
  • Nonresponse 无回应: selected people do not answer.
  • Response bias 回应偏差: people answer inaccurately (bad wording, sensitive topics).

Bias is about a consistent error in one direction – increasing the sample size does not fix it.

한국어

**편향(bias)**은 추정이 진리에 체계적으로 어긋나게 만듭니다:

  • 불완전 커버리지(undercoverage): 일부 집단이 표본 프레임에서 제외됩니다.
  • 비응답(nonresponse): 선정된 사람들이 응답하지 않습니다.
  • 응답 편향(response bias): 사람들이 부정확하게 답합니다 (나쁜 문구, 민감한 주제).

편향은 일관성 있는 한쪽 방향의 오차입니다 – 표본 크기를 늘려도 해결되지 않습니다.

편의 표본은 모집단을 누락함: 선택이 무작위가 아닐 때 편향이 유입됨
편의 표본은 모집단을 누락함: 선택이 무작위가 아닐 때 편향이 유입됨
Vocabulary · ⁨어휘⁩ Train · ⁨연습하기⁩
English 한국어
Undercoverage/ˌʌndəˈkʌvərɪdʒ/ 과소포함
Nonresponse/ˌnɒnrɪˈspɒns/ 응답 부재
Response bias/rɪˈspɒns ˈbaɪəs/ 응답 편향
3.5

Designing an Experiment · ⁨실험 설계하기⁩

Syllabus
English

Enduring Understanding (VAR-3): Well-designed experiments can establish evidence of causal relationships.

Learning Objective VAR-3.A: Identify the components of an experiment. [Skill 1.C]

  • VAR-3.A.1 The experimental units are the individuals (which may be people or other objects of study) that are assigned treatments. When experimental units consist of people, they are sometimes referred to as participants or subjects.
  • VAR-3.A.2 An explanatory variable (or factor) in an experiment is a variable whose levels are manipulated intentionally. The levels or combination of levels of the explanatory variable(s) are called treatments.
  • VAR-3.A.3 A response variable in an experiment is an outcome from the experimental units that is measured after the treatments have been administered.
  • VAR-3.A.4 A confounding variable in an experiment is a variable that is related to the explanatory variable and influences the response variable and may create a false perception of association between the two.

Learning Objective VAR-3.B: Describe elements of a well-designed experiment. [Skill 1.B]

  • VAR-3.B.1 A well-designed experiment should include the following:
    • a. Comparisons of at least two treatment groups, one of which could be a control group.
    • b. Random assignment/allocation of treatments to experimental units.
    • c. Replication (more than one experimental unit in each treatment group).
    • d. Control of potential confounding variables where appropriate.

Learning Objective VAR-3.C: Compare experimental designs and methods. [Skill 1.C]

  • VAR-3.C.1 In a completely randomized design, treatments are assigned to experimental units completely at random. Random assignment tends to balance the effects of uncontrolled (confounding) variables so that differences in responses can be attributed to the treatments.
  • VAR-3.C.2 Methods for randomly assigning treatments to experimental units in a completely randomized design include using a random number generator, a table of random values, drawing chips without replacement, etc.
  • VAR-3.C.3 In a single-blind experiment, subjects do not know which treatment they are receiving, but members of the research team do, or vice versa.
  • VAR-3.C.4 In a double-blind experiment neither the subjects nor the members of the research team who interact with them know which treatment a subject is receiving.
  • VAR-3.C.5 A control group is a collection of experimental units either not given a treatment of interest or given a treatment with an inactive substance (placebo) in order to determine if the treatment of interest has an effect.
  • VAR-3.C.6 The placebo effect occurs when experimental units have a response to a placebo.
  • VAR-3.C.7 For randomized complete block designs, treatments are assigned completely at random within each block.
  • VAR-3.C.8 Blocking ensures that at the beginning of the experiment the units within each block are similar to each other with respect to at least one blocking variable. A randomized block design helps to separate natural variability from differences due to the blocking variable.
  • VAR-3.C.9 A matched pairs design is a special case of a randomized block design. Using a blocking variable, subjects (whether they are people or not) are arranged in pairs matched on relevant factors. Matched pairs may be formed naturally or by the experimenter. Every pair receives both treatments by randomly assigning one treatment to one member of the pair and subsequently assigning the remaining treatment to the second member of the pair. Alternately, each subject may get both treatments.
한국어

지속적 이해(VAR-3): 잘 설계된 실험은 인과관계(causal relationships)에 대한 증거를確立할 수 있다.

학습 목표 VAR-3.A: 실험의 구성 요소를 식별할 수 있다. [Skill 1.C]

  • VAR-3.A.1 실험 단위(experimental units)는 treatment을 배정받는 개인들(사람 또는 기타 연구 대상 물체)이다. 실험 단위가 사람일 경우, 이들은 때때로 참여자(participants) 또는 피험자(subjects)라고 불리기도 한다.
  • VAR-3.A.2 실험에서의 설명 변수(explanatory variable) 또는 요인(factor)은 수준(levels)이 의도적으로 조작되는 변수이다. 설명 변수의 수준 또는 수준의 조합은 treatment이라고 부른다.
  • VAR-3.A.3 실험에서의 반응 변수(response variable)는 treatment이 적용된 후에 실험 단위로 부터 측정되는 결과(output)이다.
  • VAR-3.A.4 실험에서의 교란 변수(confounding variable)는 설명 변수와 관련이 있으며 반응 변수에 영향을 미칠 수 있어 두 변수 사이에 가짜 연관성을 creation할 수 있는 변수이다.

학습 목표 VAR-3.B: 잘 설계된 experiment의 요소들을 설명할 수 있다. [Skill 1.B]

  • VAR-3.B.1 잘 설계된 experiment는 다음을 포함해야 한다:
    • a. 최소 두 개의 treatment 그룹 간의 비교, 그 중 하나는 control group일 수 있다.
    • b. 실험 단위에 대한 treatment의 무작위 배정/allocation.
    • c. 반복(replication)(각 treatment 그룹당 하나 이상의 실험 단위).
    • d. 필요한 경우潜在的 confounding variables에 대한 통제(control).

학습 목표 VAR-3.C: 실험 설계와 방법을 비교할 수 있다. [Skill 1.C]

  • VAR-3.C.1 완전 무작위 설계(completely randomized design)에서는 treatments가 실험 단위에 완전히 무작위로 배정된다. 무작위 배정은 통제되지 않은(교란) 변수의 효과를 balancing하는 경향이 있어, 반응의 차이를 treatments로 귀결할 수 있게 한다.
  • VAR-3.C.2 완전 무작위 설계에서 treatments를 실험 단위에 무작위로 배정하는 방법으로는 무작위 번호 생성기(random number generator) 사용, 무작위 값 표(table of random values) 사용, 치프(chips)를 교체 없이 뽑기 등 있다.
  • VAR-3.C.3 단일 맹검 실험(single-blind experiment)에서는 피험자들이 어떤 treatment을 받고 있는지 모르고, 연구진(R&D team)은 알거나, 반대로 연구진은 모르고 피험자는 아는 상태이다.
  • VAR-3.C.4 이중 맹검 실험(double-blind experiment)에서는 피험자와 피험자를 interacts하는 연구진 구성원 모두 피험자가 어떤 treatment을 받고 있는지 모른다.
  • VAR-3.C.5 control group은 관심 있는 treatment의 효과가 있는지 확인하기 위해,感兴趣的treatment을 받지 않거나 무효 물질(placebo)이 포함된 treatment을 받는 실험 단위의 집합이다.
  • VAR-3.C.6 위약 효과(placebo effect)는 실험 단위가 placebo에 반응할 때 발생한다.
  • VAR-3.C.7 무작위 완전 블록 설계(randomized complete block designs)에서는 treatments가 각 블록(block) 내에서 완전히 무작위로 배정된다.
  • VAR-3.C.8 blocking은 실험 시작 시점에 각 블록 내부의 단위들이 적어도 하나의 blocking 변수에 대해 서로 유사하도록 보장한다. 무작위 블록 설계는 자연적 변이(natural variability)와 blocking 변수에 의한 차이(differences due to the blocking variable)를 분리하는 데 도움이 된다.
  • VAR-3.C.9 매칭된 쌍 설계는 무작위 블록 설계의 특수한 경우입니다. 차단 변수를 사용하여 연구 대상(사람일 수도 있고 그렇지 않을 수도 있음)이 관련 요인에 따라 쌍으로 매칭됩니다. 매칭된 쌍은 자연적으로 형성되거나 실험자가 직접 구성할 수 있습니다. 모든 쌍은 두 가지 처리 조건을 모두 받으며, 한 쌍의 구성원 중 한 명에게 무작위로 첫 번째 처리 조건을 배정하고 나머지 구성원에게는 두 번째 처리 조건을 배정합니다. 대안적으로 각 연구 대상이 두 가지 처리 조건을 모두 받을 수도 있습니다.

Source: College Board AP Course and Exam Description · ⁨출처: College Board AP Course and Exam Description⁩

English

Good experiments follow three principles:

  • Comparison with a control group 对照组 (often a placebo 安慰剂).
  • Random assignment 随机分配 of subjects to treatments, to balance out other variables.
  • Replication 重复: enough subjects per treatment to see a real effect.

Confounding 混杂 occurs when another variable is tied to the treatment so their effects cannot be separated; random assignment guards against it. Blinding 盲法 hides who is getting which treatment to prevent expectation effects: in a single-blind 单盲 study only one side is kept unaware (usually the subjects, or only the people assessing the result), while in a double-blind 双盲 study neither the subjects nor the researchers who interact with them know, blocking both the placebo effect and biased assessment. Blocking 区组 groups similar subjects and randomizes within each block to reduce variability.

한국어

좋은 실험은 세 가지 원칙을 따릅니다:

  • **대조군(control group)**과의 비교 (종종 위약 placebo임).
  • 다른 변수들을 균형 있게 만들기 위한 처리에 대한 무작위 배정(random assignment).
  • 반복(replication): 실제 효과를 확인할 수 있도록 처리당 충분한 수의 대상자.
완전 무작위 실험은 처리군과 대조군을 비교함
완전 무작위 실험은 처리군과 대조군을 비교함

**혼란(confounding)**은 다른 변수가 처리와 연관되어 있어 그 영향을 분리할 수 없을 때 발생합니다; 무작위 배정은 이를 방지합니다. **블라인딩(blinding)**은 누가 어떤 처리를 받는지를 숨겨 기대 효과를 방지합니다: 단일 블라인드(single-blind) 연구에서는 한쪽만 모르게 합니다 (보통 대상자이거나, 결과를 평가하는 사람만), 이중 블라인드(double-blind) 연구에서는 두 사람 모두 대상자와 상호작용하는 연구원이 모르게 하여 위약 효과와 편향된 평가를 모두 차단합니다. **블로킹(blocking)**은 유사한 대상자들을 그룹으로 나누고 각 블록 내에서 무작위 배정을 하여 변이를 줄입니다.

임상 시험: 무작위 배정이 처리와 대조를 분리함
임상 시험: 무작위 배정이 처리와 대조를 분리함
Vocabulary · ⁨어휘⁩ Train · ⁨연습하기⁩
English 한국어
control group/kənˈtrəʊl ɡruːp/ 통제군
placebo/pləˈsiːbəʊ/ 위약
Random assignment/ˈrændəm əˈsaɪnmənt/ 무작위 배정
Replication/ˌreplɪˈkeɪʃn/ 반복
Confounding/kənˈfaʊndɪŋ/ 교란 변수
Blinding/ˈblaɪndɪŋ/ 차명법
single-blind/ˈsɪŋɡl blaɪnd/ 단일 맹검
double-blind/ˈdʌbl blaɪnd/ 이중 맹검
Blocking/ˈblɒkɪŋ/ 블로킹(군집화)
3.6

Choosing the Right Design · ⁨적절한 설계 선택⁩

Syllabus
English

Enduring Understanding (VAR-3): Well-designed experiments can establish evidence of causal relationships.

Learning Objective VAR-3.D: Explain why a particular experimental design is appropriate. [Skill 1.C]

  • VAR-3.D.1 There are advantages and disadvantages for each experimental design depending on the question of interest, the resources available, and the nature of the experimental units.
한국어

지속적 이해(VAR-3): 잘 설계된 실험은 인과관계(causal relationships)에 대한 증거를確立할 수 있다.

학습 목표 VAR-3.D: 특정 실험 설계가 적절한 이유를 설명하시오. [기술 1.C]

  • VAR-3.D.1 각 실험 설계에는 관심 있는 질문,可利用한 자원, 실험 단위의 성격에 따라 장단점이 존재합니다.

Source: College Board AP Course and Exam Description · ⁨출처: College Board AP Course and Exam Description⁩

English

Match the design to the goal: use a completely randomized design for uniform subjects; a randomized block design when a known variable (sex, age) affects the response; a matched-pairs design when each subject can serve as its own control. State how you would carry out the randomization.

한국어

목표에 맞춰 설계를 선택하십시오: 균일한 대상자에게는 **완전 무작위 설계(completely randomized design)**를,已知된 변수(성별, 연령 등)가 반응에 영향을 미칠 때는 **무작위 블록 설계(randomized block design)**를, 각 대상자가 자신의 대조군이 될 수 있을 때는 매칭 페어즈(matched-pairs) 설계를 사용하십시오. 무작위 배정을 어떻게 수행할지 명시하십시오.

3.7

What an Experiment Lets You Conclude · ⁨실험으로 도출할 수 있는 결론⁩

Syllabus
English

Enduring Understanding (VAR-3): Well-designed experiments can establish evidence of causal relationships.

Learning Objective VAR-3.E: Interpret the results of a well-designed experiment. [Skill 4.B]

  • VAR-3.E.1 Statistical inference attributes conclusions based on data to the distribution from which the data were collected.
  • VAR-3.E.2 Random assignment of treatments to experimental units allows researchers to conclude that some observed changes are so large as to be unlikely to have occurred by chance. Such changes are said to be statistically significant.
  • VAR-3.E.3 Statistically significant differences between or among experimental treatment groups are evidence that the treatments caused the effect.
  • VAR-3.E.4 If the experimental units used in an experiment are representative of some larger group of units, the results of an experiment can be generalized to the larger group. Random selection of experimental units gives a better chance that the units will be representative.
한국어

지속적 이해(VAR-3): 잘 설계된 실험은 인과관계(causal relationships)에 대한 증거를確立할 수 있다.

학습 목표 VAR-3.E: 잘 설계된 실험의 결과를 해석하시오. [기술 4.B]

  • VAR-3.E.1 통계적 추론은 수집된 데이터가 도출된 분포에 기반하여 도출된 결론을 귀인합니다.
  • VAR-3.E.2 처리 조건의 무작위 배정은 연구자가 관측된 일부 변화가 우연히 일어날 가능성이 매우 낮다는 것을 결론지을 수 있게 합니다. 이러한 변화는 통계적으로 유의미하다고 합니다.
  • VAR-3.E.3 실험 처리 그룹 간 또는 내의 통계적으로 유의미한 차이는 처리 조건이 결과에 원인이 되었다는 증거입니다.
  • VAR-3.E.4 실험에 사용된 실험 단위가 더 큰 집단의 일부를 대표한다면, 실험 결과를 그 더 큰 집단에 일반화할 수 있습니다. 실험 단위의 무작위 선택은 단위가 대표성을 갖게 될 확률을 높여줍니다.

Source: College Board AP Course and Exam Description · ⁨출처: College Board AP Course and Exam Description⁩

English

Two questions decide the scope of a conclusion:

  • Random assignment used? Then a significant difference can be attributed to the treatment (causation) – for these subjects.
  • Random sampling from a population? Then results generalize to that population.

Only an experiment with random assignment supports a cause-and-effect claim; only random sampling supports generalization. Say exactly which you have.

Worked example. Researchers randomly assign $100$ volunteers to a new drug or a placebo, and the drug group improves significantly more. Because of the random assignment, the improvement can be attributed to the drug (causation) – but because the subjects were not randomly sampled, the conclusion applies only to these volunteers and does not automatically generalize to everyone.

한국어

결론의 범위를 결정하는 두 가지 질문이 있습니다:

  • 무작위 배정을 사용했는가? 그렇다면 유의미한 차이는 처리에 기인할 수 있다 (인과관계 causation) – 이 대상자에 한하여.
  • 모집단으로부터 무작위 표본을 추출했는가? 그렇다면 결과가 **해당 모집단에 일반화(generalize)**됩니다.

무작위 배정을 가진 실험만이 인과관계 주장을 지지하며, 무작위 표본만이 일반화를 지지합니다. 당신이 정확히 무엇을 가졌는지 명시하십시오.

해설 예시. 연구진들이 $100$ **지원자(volunteers)**를 새로운 약물이나 위약에 무작위로 배정하고, 약물 군이 유의미하게 더 개선되었습니다. 무작위 배정 때문에 개선은 약물에 기인할 수 있습니다 (인과관계) – 하지만 대상자들이 무작위 표본이 아니었으므로, 결론은 이 지원자들에게만 적용되며 자동으로 모든 사람에게 일반화되지는 않습니다.

3.7

Exam tips · ⁨시험 팁⁩

English
  • Distinguish an observational study (finds association) from an experiment (can show causation).
  • Good sampling is random (SRS, stratified, cluster) — beware bias (voluntary response, undercoverage, nonresponse).
  • Good experiments use control, randomization, and replication; blocking handles a known nuisance variable.
  • Only a randomized experiment supports a cause-and-effect conclusion.
  • Name the population, sample, and any confounding clearly.
한국어
  • 관찰 연구(observational study,_passive observation) (상관성 발견)와 실험(experiment, active intervention) (인과관계 입증 가능)을 구분하십시오.
  • 좋은 표본은 **무작위(SRS, 층화, 클러스터)**여야 합니다 – 편향(자발적 응답, 불완전 커버리지, 비응답)에 주의하십시오.
  • 좋은 실험은 대조군, 무작위 배정, 반복을 사용합니다; 블로킹은 알려진 번거로운 변수를 처리합니다.
  • 무작위 실험만이 인과관계 결론을 지지합니다.
  • 모집단, 표본, 그리고 혼란 요인을 명확히 명명하십시오.

Interactive lessons on this topic · ⁨이 주제에 대한 인터랙티브 수업⁩

Work through it step by step, with instant-check exercises. · ⁨즉시 체크 기능 exercises를 통해 단계별로 진행하세요.⁩

Past Papers · ⁨과거 시험지⁩

More topics in AP Statistics · ⁨AP 통계학⁩ · ⁨AP Statistics · ⁨AP 통계학⁩ 내 추가 주제⁩

Log in or create account · ⁨로그인 또는 계정 만들기⁩

IGCSE, A-Level & AP