Run-length encoding (compression) · การเข้ารหัสแบบความยาวซ้ำ (การบีบอัด)
Squeezing repeated data
- Compression makes data smaller. Run-length encoding (RLE) is one simple way.
- It works well when the same value repeats many times in a row.
- It is lossless: from the squeezed form you can rebuild the exact original.
การบีบอัดข้อมูลซ้ำซ้อน
- Compression ทำให้ข้อมูลเล็กลง Run-length encoding (RLE) เป็นวิธีง่ายๆ หนึ่งวิธี
- มันทำงานได้ดีเมื่อมีค่าเดียวกันซ้ำกันหลายครั้งติดต่อกัน
- มันคือ lossless: จากรูปแบบที่บีบอัดแล้วคุณสามารถสร้างต้นฉบับเดิมออกมาได้พอดี
What a "run" is
- A run is a stretch of the same character repeated: in
"aaabbc","aaa"is a run of 3. - RLE replaces each run with the character followed by how many times it repeats.
- So
"aaabbc"becomes"a3b2c1"— much shorter when runs are long.
"run" คืออะไร
- Run คือช่วงของ character เดียวกันซ้ำ: ใน
"aaabbc","aaa"เป็น run จำนวน 3 ครั้ง - RLE แทนที่แต่ละ run ด้วย character ตามด้วย จำนวนครั้งที่ซ้ำ
- ดังนั้น
"aaabbc"จะกลายเป็น"a3b2c1"— สั้นลงมากเมื่อ runs ยาว
Counting a run
- To measure a run, look at a character, then count how many of the same character follow it.
- Stop when the next character is different, or you reach the end (
'\0'). - That count is the run's length.
นับ run
- ในการวัด run ให้มองที่ character ก่อน แล้วนับว่ามี same character ตามมาอีกกี่ตัว
- หยุดเมื่อ character ถัดไปต่างกัน หรือคุณถึงจุดจบ (
'\0') - ค่านี้คือความยาวของ run
Building the encoded string
- Walk the input. For each run, write the character, then its count, into the output.
- Move your input position past the whole run before starting the next one.
- End the output string with
'\0'so it is a proper C string.
สร้าง string ที่เข้ารหัส
- เดินผ่าน input สำหรับแต่ละ run เขียน character, แล้วตามด้วย count, ลงใน output
- เคลื่อนตำแหน่ง input ของคุณ พ้น ทั้ง run ก่อนเริ่ม run ถัดไป
- ทิ้ง output string ด้วย
'\0'เพื่อให้เป็น C string ที่ถูกต้อง
#include <stdio.h>
int main(void) {
const char *s = "aaab";
int i = 0;
char c = s[i];
int count = 0;
while (s[i] == c) { // count the first run
count++;
i++;
}
printf("%c%d\n", c, count); // a3
return 0;
}
Common mistakes
- Run-length encoding stores a value then its count; it only helps when there are long runs.
- It is lossless — the original is restored exactly.
ข้อผิดพลาดที่พบบ่อย
- Run-length encoding เก็บค่าแล้วตามด้วย count; มันมีประโยชน์ก็ต่อเมื่อมี long runs เท่านั้น
- มัน lossless — ต้นฉบับจะถูกคืนกลับมาอย่างสมบูรณ์
Now you try
- Find each run, then write the character and its count to the output.
- The caller gives you an output buffer big enough to hold the result. Do not write a
main.
ลองดูเลย
- หาแต่ละ run แล้วเขียน character และ count ของมันลงใน output
- ผู้เรียกมอบ output buffer ที่ใหญ่พอที่จะเก็บผลลัพธ์ อย่า เขียน
main
Run-length encoding
Replace a run of repeats with count + symbol — lossless. · แทนกลุ่มซ้ำด้วย จำนวน + สัญลักษณ์ — ไม่สูญเสียข้อมูล
Complete int run_length_at(const char *s, int i) so it returns how many times the character s[i] repeats starting at index i. Do not write a main. · เติม int run_length_at(const char *s, int i) เพื่อให้คืนจำนวนครั้งที่ตัวอักษร s[i] ซ้ำกันต่อเนื่องตั้งแต่ดัชนี i ห้าม เขียน main
Click Run to see the output here. · คลิก Run เพื่อดูผลลัพธ์ที่นี่
Complete void rle_encode(const char *in, char *out) so it writes the run-length encoding of in into out: each run becomes the character then its count. "aaabbc" becomes "a3b2c1". End out with '\0'. Do not write a main. · เติม void rle_encode(const char *in, char *out) เพื่อเขียนการเข้ารหัสแบบความยาวซ้ำของ in ลงใน out: แต่ละกลุ่ม会变成สัญลักษณ์แล้วตามด้วยจำนวน "aaabbc" จะกลายเป็น "a3b2c1" จบ out ด้วย '\0' ห้าม เขียน main
Click Run to see the output here. · คลิก Run เพื่อดูผลลัพธ์ที่นี่