AlanHQ
@AlanPhan04
About
Data Engineer building GCP-based Lakehouse infra (Spark, Iceberg, Kafka, Airflow) | Final-year HCMUT CS student
Skills & Technologies
Projects & Repositories
Recent public projects and repositories from this profile.
spark-101-to-pro
TeX📚 Personal learning journey through Apache Spark — from fundamentals to advanced query optimization, with a focus on SparkSQL performance tuning and Apache Iceberg on a production Lakehouse (GCS + Hive Metastore + Kyuubi). Notes, EXPLAIN plan breakdowns, and real-world tuning case studies documented in Markdown.
customer-data-platform
PythonLive Wikipedia edits streamed through Kafka → Flink → ClickHouse, with a Hive/RustFS data lake, a Kafka-audited metastore, and a BigQuery-style query editor — one Docker Compose stack, no fake data.
AlanPhan04
No project description available.
awesome-public-datasets
A topic-centric list of HQ open datasets.
alanphan04.github.io
HTMLNo project description available.
Multidisciplinary-Project
JavaA research project in HCMUT on compression algorithms for streaming IoT data, exploring existing techniques and their applicability. Focus on compression efficiency, performance trade-offs, and real-time data processing. Future work includes real-world implementation and optimization for IoT applications.