Data compression · דחיסת נתונים
Why compress data
- Data takes up space to store and time to send.
- Compression makes a file smaller so it is cheaper to save and faster to share.
- There are two kinds: lossless and lossy.
מדוע לדחס נתונים
- נתונים תופסים מרחב אחסון וזמן לשליחה.
- דחיפה הופכת קובץ לקטן יותר כך שהוא זול יותר לאחסון ומהיר יותר להעברה.
- ישנם שני סוגים: ללא אובדן (Lossless) ו-עם אובדן (Lossy).
Lossless compression
- Lossless makes a file smaller but keeps every bit of the data.
- When you open the file, you get back the exact original.
- It works by finding patterns and writing them in a shorter way.
דחיפה ללא אובדן
- ללא איבוד מקטין קובץ אך שומר על כל הנתונים.
- כאשר פותחים את הקובץ, מתקבלים בחזרה את המקורי המדויק.
- זה עובד על ידי מציאת תבניות וכתיבה שלהן בצורה קצרה יותר.
Original: AAAAAAAA-BBB
Shorter: 8A-3B (8 A's, then 3 B's)
Open it: AAAAAAAA-BBB (exactly the same again)
Lossy compression
- Lossy makes a file much smaller by throwing away some data.
- It drops details that people can barely see or hear.
- You cannot get the exact original back — but it is "close enough".
דחיסה עם איבוד
- עם איבוד מקטין משמעותית את הקובץ על ידי ויתור על חלק מהנתונים.
- הוא מוותר על פרטים שאנשים קשה לראות או לשמוע.
- אין אפשרות לחזור למקורי המדויק — אך התוצאה "קרובה מספיק".
Photo (large) --lossy--> Photo (small)
A few colors and fine details are gone,
but your eye hardly notices.
The trade-off: size vs quality
- Lossless keeps full quality, but the file stays larger.
- Lossy gives a much smaller file, but quality goes down a little.
- You choose based on what matters more: perfect data or small size.
הפיצול: גודל מול איכות
- ללא איבוד שומר על איכות מלאה, אך הקובש נשאר גדול יותר.
- עם איבוד מספק קובץ הרבה קטן יותר, אך האיכות יורדת מעט.
- הבחירה נעשית לפי מה חשוב יותר: נתונים מושלמים או גודל קטן.
When to use each
- Use lossless when every detail must be exact.
- Use lossy when a small drop in quality is fine and small size matters.
מתי להשתמש בכל סוג
- השתמש ב-ללא איבוד כאשר כל פרט חייב להיות מדויק.
- השתמש ב-עם איבוד כאשר ירידה זניחה באיכות היא מקובלת וגודל קטן הוא קריטי.
Lossless: text, code, a .zip file, a spreadsheet
Lossy: photos (JPEG), music (MP3), video
Key idea
- Compression trades size against quality (or against work to undo it).
- Lossless = smaller and perfect; lossy = much smaller but not exact.
- Good engineers pick the right kind for the job.
הרעיון המרכזי
- דחיסה מחליפה גודל מול איכות (או מול העבודה הנדרשת להחזרתו).
- ללא איבוד = קטן יותר ומושלם; עם איבוד = הרבה קטן יותר אך לא מדויק.
- מהנדסים טובים בוחרים את הסוג המתאים למשימה.
Run-length encoding
- Run-length encoding (RLE) is a simple lossless method.
- A run is a stretch of the same character repeated. RLE stores a count instead of the repeats.
- We will store each run as a pair
[character, count]inside a list.
ריצות-לונג'ית אנקודינג
- רצף-אנקודינג (RLE) היא שיטה פשוטה ללא איבוד.
- רצף הוא רצף של אות同一个重复出现的字符。RLE 存储的是计数而不是重复的字符。
- נאחסן כל רצף בזוג
[character, count]בתוך רשימה.
"AAAB" -> [["A", 3], ["B", 1]] (3 A's, then 1 B)
[["A", 3], ["B", 1]] -> "AAAB" (decode it back — exact again)
Common mistakes
- Lossless compression can be reversed exactly; lossy throws away detail.
- More compression can mean lower quality.
טעויות נפוצות
- דחיסה ללא איבוד נתונים ניתנת להפעה במדויק; דחיסה עם איבוד נתונים מוותרת על פרטים.
- דחיסה גדולה יותר עלולה להוביל לאיכות נמוכה יותר.
Now you try
- Build RLE yourself: an encoder, a decoder, and a length helper.
- Each task checks your function on several inputs. Press Check answer.
כעת תנסו בעצמכם
- בנה RLE בעצמך: מכונן, מפענח ותומך אורך.
- כל משימה תבדוק את הפונקציה שלך מספר פעמים. לחץ על בדיקת תשובה.
Lossless compression · דחיסה ללא אובדן
Run-length encoding replaces a run of repeats with count + symbol. · קידוד אורך-רצף מחליף רצף של חזרות עם ספירה + סמל.
Write encode(text) for run-length encoding. Return a list of [character, count] pairs, one per run of repeats. Example: encode("AAAB") → [['A', 3], ['B', 1]]. For the empty string return []. · כתוב encode(text) לקידוד אורך-רצף. החזר רשימת [character, count] זוגות, אחד לכל רצף חזרות. דוגמה: encode("AAAB") → [['A', 3], ['B', 1]]. עבור מחרוזת ריקה החזר [].
Click Run to see the output here. · לחץ על הרץ כדי לראות את התוצא כאן.
Write decode(pairs) that reverses the encoder: given a list of [character, count] pairs, rebuild the original string. Example: decode([['A', 3], ['B', 1]]) → 'AAAB'. For [] return ''. · כתוב decode(pairs) שמחליף את המקודד: קבל רשימת [character, count] זוגות, ובנה מחדש את המחרוזת המקורית. דוגמה: decode([['A', 3], ['B', 1]]) → 'AAAB'. עבור [] החזר ''.
Click Run to see the output here. · לחץ על הרץ כדי לראות את התוצא כאן.
Without decoding, write original_length(pairs) that returns how many characters the original text had — just add up the counts. Example: original_length([['A', 3], ['B', 1]]) → 4. · ללא פירוש, כתוב original_length(pairs) שהחזר כמה תווים הייתה למחרוזת המקור – פשוט סכום את הספירות. דוגמה: original_length([['A', 3], ['B', 1]]) → 4.
Click Run to see the output here. · לחץ על הרץ כדי לראות את התוצא כאן.