Skip to content
View Marcus778899's full-sized avatar
😁
Happy
😁
Happy
  • STARCO TECHNOLOGY CO., LTD.
  • 3 F., No. 135, Sec. 3, Minsheng E. Rd., Songshan Dist., Taipei City 105007, Taiwan (R.O.C.)
  • 16:22 (UTC +08:00)
  • LinkedIn in/marcus-807788284

Block or report Marcus778899

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Marcus778899/README.md

Marcus Lin (Sheng-Lun Lin)

Big Data Engineer | Digital Transformation Architect

I build data platforms β€” consolidating data scattered across systems into a layered, governed, reusable asset, so business teams never have to ask "where does this number come from" twice.

My background spans data engineering and backend development; my previous role was backend full-time. That means I don't just consume data downstream β€” I can intervene at the source: system design, API contracts, and the moment an event is produced. In platform work, that is often what decides success or failure.

With 3+ years of experience, I care less about which tools are on the list and more about designing the right thing for the situation.


What I Handle

Building a Data Platform from Zero

The company has data but no architecture; reports are assembled by hand and metric definitions differ across departments. I start from raw landing in the data lake and design the full layering β€” staging, ODS, data warehouse, data marts β€” defining the responsibility and metric semantics of each layer so data is traceable, re-runnable, and shareable across business units.

Data Modeling and Metric Governance

I use a bus matrix to align business processes with conformed dimensions β€” clarifying who needs to see what, and at what grain, before deciding on fact and dimension design. This avoids the common trap where every department builds its own model and the numbers stop agreeing.

Pipeline Design and Operations

I orchestrate the full data lifecycle with Airflow: scheduling dependencies, incremental versus full-load strategies, retry semantics, and data quality checks. The goal is not a pipeline that runs β€” it is a pipeline where, when something breaks, you know exactly where it broke and can recover safely.

Real-Time and Change Data Capture

I use CDC (Debezium) with Kafka to capture change events from source systems, moving reporting from T+1 batch toward near real-time without adding query pressure to the source database.

Analytical Performance Tuning

Against terabytes of historical data, I apply OLAP engines with partitioning and indexing strategies to bring multi-minute queries down to seconds β€” turning analysis from waiting in line for results into live exploration.

Backend Services and Data Interfaces

I design and implement data service APIs that deliver platform output to frontends, reporting layers, and external systems. I also work in the other direction, on event design and write paths in source systems. The wall between backend and data engineering is one I can stand on both sides of.

Data Access Layer for AI Agents

I build MCP servers so AI agents can access internal enterprise data under controlled permissions and boundaries β€” the bridge between data governance and AI applications.


Technical Scope

Category Technologies
Cloud GCP, AWS
Storage & Query BigQuery, ClickHouse, Redshift, PostgreSQL, MongoDB, Redis
Streaming & CDC Kafka, Debezium, Spark
Transformation & Orchestration dbt, Airflow
Backend & Deployment FastAPI, Spring Boot, Docker, Docker Compose
AI MCP Server, LangGraph

Technical Interests

Database internals and distributed systems. Currently exploring advanced data modeling techniques and improving observability across large-scale data infrastructure.


Connect

Pinned Loading

  1. resume resume Public

    resume

    HTML

  2. finance_project finance_project Public

    Python

  3. python_tutorial python_tutorial Public

    Jupyter Notebook

  4. Ecpay_api Ecpay_api Public

    Python

  5. chatbot chatbot Public

    Python