Run-length encoding (compression) · קידוד אורך רצף (דחיסה)
Squeezing repeated data
- Compression makes data smaller. Run-length encoding (RLE) is one simple way.
- It works well when the same value repeats many times in a row.
- It is lossless: from the squeezed form you can rebuild the exact original.
צפיפת נתונים חוזרים
- צפיפות הופכת נתונים לקטנים יותר. העתקה באורך ריצה (RLE) היא אחת הדרכים הפשוטות.
- היא פועלת היטב כאשר אותה ערך חוזר פעמים רבות ברצף.
- היא ללא איבוד: מהצורה הצפופה ניתן לשחזר בדיוק את המקור.
What a "run" is
- A run is a stretch of the same character repeated: in
"aaabbc","aaa"is a run of 3. - RLE replaces each run with the character followed by how many times it repeats.
- So
"aaabbc"becomes"a3b2c1"— much shorter when runs are long.
מהי "ריצה"
- ריצה היא רצף של אותו קריטריון חוזר: ב
"aaabbc","aaa"היא ריצה של 3. - RLE מחליפה כל ריצה באות שלאחריו כמה פעמים הוא חוזר.
- כך
"aaabbc"הופך ל"a3b2c1"— קצר הרבה כשהריצות ארוכות.
Counting a run
- To measure a run, look at a character, then count how many of the same character follow it.
- Stop when the next character is different, or you reach the end (
'\0'). - That count is the run's length.
ספירת ריצה
- כדי למדוד ריצה, הסתכלו על אות, ואז ספירו כמה מאות אותו האות נערכים אחריו.
- עצרו כאשר האות הבא שונה, או שהגעתם לסוף (
'\0'). - מספר זה הוא אורך הריצה.
Building the encoded string
- Walk the input. For each run, write the character, then its count, into the output.
- Move your input position past the whole run before starting the next one.
- End the output string with
'\0'so it is a proper C string.
בניית הממחושק
- עבר על הקלט. עבור כל רצף, כתוב את התווית, ולאחר מכן את הספירה, לתוצאה.
- הזז את נקודת הקלט שלך מעבר לרצף כולו לפני שמתחילים ברצף הבא.
- סיים את מחרוזת התוצאה עם
'\0'כדי שתהיה מחרוזת C תקינה.
#include <stdio.h>
int main(void) {
const char *s = "aaab";
int i = 0;
char c = s[i];
int count = 0;
while (s[i] == c) { // count the first run
count++;
i++;
}
printf("%c%d\n", c, count); // a3
return 0;
}
Common mistakes
- Run-length encoding stores a value then its count; it only helps when there are long runs.
- It is lossless — the original is restored exactly.
טעויות נפוצות
- ממחושק אורך-רצף (Run-length encoding) מאחסן ערך ולאחר מכן את הספירה שלו; הוא מועיל רק כאשר ישנם רצפים ארוכים.
- זהו קידוד לא אבוד — המקור משוחזר בדיוק.
Now you try
- Find each run, then write the character and its count to the output.
- The caller gives you an output buffer big enough to hold the result. Do not write a
main.
כעת תנסו בעצמכם
- מצא כל רצף, ולאחר מכן כתוב את התווית ואת הספירה שלה לתוצאה.
- הקורא מספק לך буפר תוצאה גדול מספיק כדי להכיל את התוצאה. אל תכתוב
main.
Run-length encoding · אנכוד רצפים
Replace a run of repeats with count + symbol — lossless. · החלף רצף חזרות ב-מספר + סימן — ללא איבוד נתונים.
Complete int run_length_at(const char *s, int i) so it returns how many times the character s[i] repeats starting at index i. Do not write a main. · שלים int run_length_at(const char *s, int i) כך שיחזיר כמה פעמים התווית s[i] חוזרת החל מאינדקס i. אל תכתוב main.
Click Run to see the output here. · לחץ על הרץ כדי לראות את התוצא כאן.
Complete void rle_encode(const char *in, char *out) so it writes the run-length encoding of in into out: each run becomes the character then its count. "aaabbc" becomes "a3b2c1". End out with '\0'. Do not write a main. · שלים void rle_encode(const char *in, char *out) כדי שייכתוב את קידוד אורך הרצף של in לתוך out: כל רצף הופך לתווית ולאחריה המספר שלה. "aaabbc" הופך ל-"a3b2c1". סיים out עם '\0'. אל תכתוב main.
Click Run to see the output here. · לחץ על הרץ כדי לראות את התוצא כאן.