Run-length encoding (compression)
This page needs a recent browser (with SharedArrayBuffer support). Please update Chrome, Edge, Firefox or Safari to the latest version. · 이 페이지는 최신 브라우저(SharedArrayBuffer 지원)가 필요합니다. Chrome, Edge, Firefox 또는 Safari를 최신 버전으로 업데이트해 주세요.
English
Squeezing repeated data
- Compression makes data smaller. Run-length encoding (RLE) is one simple way.
- It works well when the same value repeats many times in a row.
- It is lossless: from the squeezed form you can rebuild the exact original.
한국어
반복되는 데이터 압축하기
- 압축(compression) 은 데이터를 작게 만듭니다. Run-length encoding (RLE) 은 그중 하나인 간단한 방법입니다.
- 같은 값이 연속해서 여러 번 나타날 때 잘 작동합니다.
- 무손실(lossless) 입니다: 압축된 형태로부터 정확히 원본을 복원할 수 있습니다.
English
What a "run" is
- A run is a stretch of the same character repeated: in
"aaabbc","aaa"is a run of 3. - RLE replaces each run with the character followed by how many times it repeats.
- So
"aaabbc"becomes"a3b2c1"— much shorter when runs are long.
한국어
'run'이란 무엇인가
- run은 동일한 문자가 연속해서 반복되는 부분입니다:
"aaabbc"에서"aaa"는 3번 반복된 run입니다. - RLE은 각 run을 '문자 + 반복 횟수'로 교체합니다.
- 따라서
"aaabbc"는"a3b2c1"로 변하며, run이 길면 훨씬 짧아집니다.
English
Counting a run
- To measure a run, look at a character, then count how many of the same character follow it.
- Stop when the next character is different, or you reach the end (
'\0'). - That count is the run's length.
한국어
run 세기
- run의 길이를 측정하려면 한 문자를 보고, 그 뒤에 이어지는 같은 문자가 몇 개인지 세십시오.
- 다음 문자가 다르거나 끝(
'\0')에 도달할 때까지 세기를 멈춥니다. - 그 세기가 바로 run의 길이입니다.
English
Building the encoded string
- Walk the input. For each run, write the character, then its count, into the output.
- Move your input position past the whole run before starting the next one.
- End the output string with
'\0'so it is a proper C string.
한국어
인코딩된 문자열 만들기
- 입력을 순회하십시오. 각 run마다 문자와 계수(count) 를 출력에 기록하십시오.
- 다음 run을 시작하기 전에 입력 위치를 전체 run 바깥으로 이동시키십시오.
- C 문자열로 올바르게 인식되도록 출력 문자열을
'\0'로 끝내십시오.
#include <stdio.h>
int main(void) {
const char *s = "aaab";
int i = 0;
char c = s[i];
int count = 0;
while (s[i] == c) { // count the first run
count++;
i++;
}
printf("%c%d\n", c, count); // a3
return 0;
}
English
Common mistakes
- Run-length encoding stores a value then its count; it only helps when there are long runs.
- It is lossless — the original is restored exactly.
한국어
흔한 실수
- Run-length encoding은 값과 계수를 저장하므로, 긴 run이 있을 때만 도움이 됩니다.
- 무손실(lossless)입니다 — 원본이 정확히 복원됩니다.
English
Now you try
- Find each run, then write the character and its count to the output.
- The caller gives you an output buffer big enough to hold the result. Do not write a
main.
한국어
이제 직접 해보기
- 각 run을 찾아서 문자와 그 계수를 출력에 기록하십시오.
- caller가 결과를 담을 수 있는 충분한 크기의 출력 버퍼를 제공해 줍니다. 반드시
main를 작성하지 마십시오.
Explore · 탐색하기
Run-length encoding
Replace a run of repeats with count + symbol — lossless.
Complete int run_length_at(const char *s, int i) so it returns how many times the character s[i] repeats starting at index i. Do not write a main.
Click Run to see the output here. · 출력을 보려면 '실행'을 클릭하세요.
Complete void rle_encode(const char *in, char *out) so it writes the run-length encoding of in into out: each run becomes the character then its count. "aaabbc" becomes "a3b2c1". End out with '\0'. Do not write a main.
Click Run to see the output here. · 출력을 보려면 '실행'을 클릭하세요.