บทที่ 1: Data Engineering Described — Data Engineering คืออะไร?
4 min readคำนิยามจาก Experts ต่างๆ
- AlexSoft: ชุด operations ที่สร้าง interfaces และ mechanisms สำหรับ flow และการเข้าถึงข้อมูล
- Jesse Anderson: แบ่งเป็น 2 แบบ — SQL-focused (ใช้ RDBMS — Relational Database Management System + SQL/ETL — Extract, Transform, Load) กับ Big Data–focused (ใช้ Hadoop, Spark, Flink)
- Maxime Beauchemin (creator of Airflow): superset ของ BI และ data warehousing ที่ดึงองค์ประกอบจาก software engineering เข้ามา
- Lewis Gavin: everything about movement, manipulation, และ management ของ data
นิยามของหนังสือเล่มนี้: Data engineering คือการ develop, implement, และ maintain systems/processes ที่รับ raw data เข้ามาแล้วผลิต high-quality, consistent information เพื่อรองรับ downstream use cases (analysis, ML) — เป็นจุดตัดของ security, data management, DataOps, data architecture, orchestration และ software engineering ทั้งหมดนี้ถูกสรุปเป็น Data Engineering Lifecycle (5 stages + 6 undercurrents ดูรายละเอียดในบทที่ 2)
Figure 1-1. The data engineering lifecycle — ภาพรวมทั้งเล่มที่ทุกบทจะอ้างอิงกลับมา
วิวัฒนาการของ Data Engineering
| ยุค | จุดเด่น |
|---|---|
| 1980–2000: Data Warehousing | Bill Inmon บัญญัติ data warehouse (1989), Kimball พัฒนา dimensional modeling, MPP databases (Massively Parallel Processing — ประมวลผลแบบขนาน), ETL developers |
| 2000s: Birth of Big Data | Google ตีพิมพ์ GFS — Google File System (2003) และ MapReduce (2004) — "big bang" ของ data engineering, Hadoop ถือกำเนิดที่ Yahoo (2006), AWS เปิดตัว |
| 2010s: Big Data Era | Hadoop ecosystem เติบโต (Hive, Pig, HBase, Spark, Presto), code-first engineering, แต่หลายองค์กรใช้ big data tools กับ data น้อยเกินไป |
| 2020s: Data Lifecycle Engineering | Modern data stack — abstraction สูงขึ้น, managed services, data engineer กลายเป็น "data lifecycle engineer" |
Skills ที่ Data Engineer ต้องมี
Technical:
- SQL — Structured Query Language: lingua franca ของ data, กลับมาสำคัญอีกครั้ง
- Python — bridge language ระหว่าง data engineering และ data science
- JVM languages (Java/Scala) — พบใน Apache open source projects
- bash — command-line scripting ใน data pipeline
- เข้าใจ software engineering, networking, distributed computing, storage
Business:
- Communication กับทั้ง technical และ non-technical people
- Requirement scoping — รู้ว่าควร build อะไรและผลกระทบต่อธุรกิจ
- Cost control — optimize time to value, TCO (Total Cost of Ownership — ต้นทุนรวมในการเป็นเจ้าของ), opportunity cost
- เข้าใจ Agile, DevOps, DataOps culture
ความแตกต่างของบทบาท
| บทบาท | โฟกัส |
|---|---|
| Data Engineer | สร้าง foundation — get data, store, process, prepare |
| Data Scientist | สร้าง forward-looking models, predictions, recommendations |
| Data Analyst | โฟกัส past/present, ใช้ SQL, BI tools (Business Intelligence), domain experts |
| ML Engineer | คร่อม data engineering และ data science, train models ใน production |
| AI Researcher | Advanced ML techniques (GPT, DALL-E), มักอยู่ใน big tech/academia |
Figure 1-5. The Data Science Hierarchy of Needs (Monica Rogati, 2017) — data scientist ใช้เวลา 70-80% วนอยู่ 3 ชั้นล่าง (gather/clean/process data) ก่อนจะถึงชั้น AI/ML ด้านบน จึงต้องมี data engineer สร้าง foundation ให้ก่อน
ประเภทของ Data Engineers
- Type A (Abstraction): ใช้ off-the-shelf products, managed services, ทำให้ architecture simple ที่สุด ไม่ reinvent the wheel — พบได้ทุก data maturity stage
- Type B (Build): Build custom tools และ systems ที่สร้าง competitive advantage — มักพบใน stage 2-3 (scaling/leading) หรือเมื่อ use case unique จนต้องสร้างเอง
- ในทางปฏิบัติ Type A/B มักอยู่ในทีมเดียวกัน (หรือเป็นคนเดียวกัน) — Type A มักถูกจ้างมาก่อนเพื่อวาง foundation แล้ว Type B skillset ค่อยตามมาทีหลังตามความจำเป็น
Internal-Facing vs External-Facing Data Engineer
Figure 1-9. The directions a data engineer faces
- External-facing: ดูแล data จาก external-facing apps (social media, IoT, ecommerce) — มี feedback loop กลับไปที่ application ต้องรับมือกับ concurrency สูงและ security ที่ซับซ้อนกว่า (โดยเฉพาะ multitenant data — ข้อมูลหลาย customer ในตารางเดียว)
- Internal-facing: โฟกัส pipeline/warehouse สำหรับ BI dashboards, reports, data science, ML ภายในองค์กร
- ในทางปฏิบัติสองบทบาทนี้มักผสมกัน — internal-facing data มักเป็น prerequisite ของ external-facing data
ระดับ Data Maturity
- Starting with Data — ทีมเล็ก, generalist, ad hoc requests, ต้อง get buy-in และ build foundation
- Scaling with Data — มี formal practices แล้ว, ย้ายจาก generalist สู่ specialist, ใช้ DevOps/DataOps
- Leading with Data — data-driven organization, self-service analytics, automation, data governance
Stakeholders ของ Data Engineer
Figure 1-12. Key technical stakeholders of data engineering
- Upstream (data producers): data architects (ออกแบบ blueprint ของ data architecture ทั้งองค์กร), software engineers (สร้าง internal data จาก application events/logs), DevOps/SREs (operational monitoring data)
- Downstream (data consumers): data scientists, data analysts, ML engineers/AI researchers (ตามตารางด้านบน)
- C-suite: CEO/CIO/CTO วางกลยุทธ์และ sponsor initiatives, ส่วน CDO (Chief Data Officer — ดูแล data asset/strategy) และ CAO (Chief Analytics/Algorithms Officer — ดูแล analytics/ML strategy) เป็น role เฉพาะทางด้าน data ที่พบมากขึ้นเรื่อยๆ
- Data engineer ยังทำงานร่วมกับ project managers (จัดการ sprint/timeline) และ product managers (ดูแล data products)