Data compression · การบีบอัดข้อมูล
Why compress data
- Data takes up space to store and time to send.
- Compression makes a file smaller so it is cheaper to save and faster to share.
- There are two kinds: lossless and lossy.
ทำไมต้องบีบอัดข้อมูล
- ข้อมูลกิน พื้นที่ ในการจัดเก็บและ เวลา ในการส่ง
- การบีบอัด ทำให้ไฟล์เล็กลงเพื่อประหยัดค่าจัดเก็บและส่งต่อได้เร็วขึ้น
- มีสองประเภท: lossless และ lossy
Lossless compression
- Lossless makes a file smaller but keeps every bit of the data.
- When you open the file, you get back the exact original.
- It works by finding patterns and writing them in a shorter way.
Lossless compression
- Lossless ทำให้ไฟล์เล็กลงแต่เก็บ ทุก บิตของข้อมูลไว้ครบถ้วน
- เมื่อคุณเปิดไฟล์กลับมา จะได้ เดิม ที่สมบูรณ์แบบ
- ทำงานโดยการหารูปแบบและเขียนด้วยวิธีที่สั้นกว่า
Original: AAAAAAAA-BBB
Shorter: 8A-3B (8 A's, then 3 B's)
Open it: AAAAAAAA-BBB (exactly the same again)
Lossy compression
- Lossy makes a file much smaller by throwing away some data.
- It drops details that people can barely see or hear.
- You cannot get the exact original back — but it is "close enough".
Lossy compression
- Lossy ทำให้ไฟล์เล็กลงมาก bằng ทิ้ง ข้อมูลบางส่วนออก
- ตัดรายละเอียดที่คนมองเห็นหรือได้ยินยาก
- คุณ ไม่สามารถ ได้ค่าเดิมที่สมบูรณ์กลับมาได้ — แต่ถือว่า "ใกล้เคียงพอใช้"
Photo (large) --lossy--> Photo (small)
A few colors and fine details are gone,
but your eye hardly notices.
The trade-off: size vs quality
- Lossless keeps full quality, but the file stays larger.
- Lossy gives a much smaller file, but quality goes down a little.
- You choose based on what matters more: perfect data or small size.
The trade-off: size vs quality
- Lossless รักษาคุณภาพเต็มรูปแบบแต่ไฟล์จะใหญ่ขึ้น
- Lossy ให้ไฟล์เล็กมากแต่คุณภาพลดลงเล็กน้อย
- เลือกตามสิ่งที่สำคัญมากกว่า: ข้อมูลสมบูรณ์ หรือ ขนาดเล็ก
When to use each
- Use lossless when every detail must be exact.
- Use lossy when a small drop in quality is fine and small size matters.
When to use each
- ใช้ lossless เมื่อต้องรักษาทุกรายละเอียดให้ถูกต้อง
- ใช้ lossy เมื่อการลดคุณภาพเล็กน้อยยอมรับได้และต้องการขนาดเล็ก
Lossless: text, code, a .zip file, a spreadsheet
Lossy: photos (JPEG), music (MP3), video
Key idea
- Compression trades size against quality (or against work to undo it).
- Lossless = smaller and perfect; lossy = much smaller but not exact.
- Good engineers pick the right kind for the job.
Key idea
- Compression trading size against quality (or against work to undo it)
- Lossless = เล็กและ สมบูรณ์; lossy = เล็กมากแต่ ไม่สมบูรณ์
- วิศวกรที่ดีเลือกประเภทที่เหมาะสมสำหรับงาน
Run-length encoding
- Run-length encoding (RLE) is a simple lossless method.
- A run is a stretch of the same character repeated. RLE stores a count instead of the repeats.
- We will store each run as a pair
[character, count]inside a list.
Run-length encoding
- Run-length encoding (RLE) เป็นวิธี lossless แบบง่าย
- A run คือชุดของตัวอักษรซ้ำกัน RLE เก็บจำนวนแทนการซ้ำ
- We will store each run as a pair
[character, count]inside a list
"AAAB" -> [["A", 3], ["B", 1]] (3 A's, then 1 B)
[["A", 3], ["B", 1]] -> "AAAB" (decode it back — exact again)
Common mistakes
- Lossless compression can be reversed exactly; lossy throws away detail.
- More compression can mean lower quality.
ข้อผิดพลาดที่พบบ่อย
- Lossless compression can be reversed exactly; lossy throws away detail
- More compression can mean lower quality
Now you try
- Build RLE yourself: an encoder, a decoder, and a length helper.
- Each task checks your function on several inputs. Press Check answer.
ลองดูเลย
- Build RLE yourself: an encoder, a decoder, and a length helper
- Each task checks your function on several inputs. Press Check answer
Lossless compression · การบีบอัดแบบไม่สูญเสียข้อมูล
Run-length encoding replaces a run of repeats with count + symbol. · Run-length encoding เปลี่ยนชุดซ้ำซ้อนด้วย จำนวน + สัญลักษณ์
Write encode(text) for run-length encoding. Return a list of [character, count] pairs, one per run of repeats. Example: encode("AAAB") → [['A', 3], ['B', 1]]. For the empty string return []. · เขียน encode(text) สำหรับ run-length encoding กลับ ลิสต์ของคู่ [character, count], คู่ละหนึ่งสำหรับแต่ละชุดซ้ำซ้อน ตัวอย่าง: encode("AAAB") → [['A', 3], ['B', 1]] สำหรับสตริงว่าง ให้กลับ []
Click Run to see the output here. · คลิก Run เพื่อดูผลลัพธ์ที่นี่
Write decode(pairs) that reverses the encoder: given a list of [character, count] pairs, rebuild the original string. Example: decode([['A', 3], ['B', 1]]) → 'AAAB'. For [] return ''. · เขียน decode(pairs) ที่กลับด้าน encoder: จากลิสต์คู่ [character, count] สร้างสตริงเดิมกลับมา ตัวอย่าง: decode([['A', 3], ['B', 1]]) → 'AAAB' สำหรับ [] ให้กลับ ''
Click Run to see the output here. · คลิก Run เพื่อดูผลลัพธ์ที่นี่
Without decoding, write original_length(pairs) that returns how many characters the original text had — just add up the counts. Example: original_length([['A', 3], ['B', 1]]) → 4. · โดยไม่มีการ decode ให้เขียน original_length(pairs) ที่กลับจำนวนตัวละครที่ข้อความเดิมมี — простоบวกจำนวนทั้งหมดเข้าด้วยกัน ตัวอย่าง: original_length([['A', 3], ['B', 1]]) → 4
Click Run to see the output here. · คลิก Run เพื่อดูผลลัพธ์ที่นี่