01 · Roasts
Pipeline, meet test suite
Log-Realtime-Analysis orchestrates Kafka, Spark, DynamoDB, Dash, and five Docker services—then ships with zero tests and zero CI.
Commit heatmap: cameo role
54 commits this year and long empty stretches on the heatmap make the 2026 push look more like a guest appearance than a season.
Notebook gravity
59% Jupyter Notebook is on-brand for an ML portfolio; turning one analysis into a maintained package would change the signal.
Adoption is still warming up
32 total stars, 8 forks, and 17 followers show real interest, but none of the three projects has evidence of users beyond GitHub reactions.
Built using
Zoral
Shadows one worker for a week, then takes over their job with zero extra setup. Behaves exactly like the original.
zoral.ai
02 · Category breakdown
- Impact25% weight33F
- Consistency20% weight35F
- Quality20% weight52D
- Depth15% weight35F
- Breadth10% weight65C
- Community10% weight40D
03 · Stats
365-day commit heatmap
53 active days
Language distribution
- Jupyter Notebook59%
- HTML28%
- TypeScript7%
- Python3%
- Blade1%
- PHP1%
- Other1%
04 · Numbers
Owned repos
non-fork
29
Commits
last 12 months
54
Followers
17
Joined GitHub
Mar 2017
05 · Top repos
ronaldkanyepi /
Log-Realtime-Analysis
A documented, Dockerized Kafka–Spark–DynamoDB–Dash pipeline with substantial integration scope and visualization, but limited by one-day history, no tests or CI, untyped Python, and little demonstrated adoption.
ronaldkanyepi /
Geolocation-App-Python-Streamlit-Folium
A small, documented Streamlit geolocation demo that combines IPGeolocation API data with an interactive Folium map, but remains a single-file application with limited engineering safeguards and adoption.
ronaldkanyepi /
Customer-Churn-Analysis
A documented, reproducible-looking telecom churn notebook analyzes 7,043 records with EDA, feature engineering, and a broad modeling stack, but remains a one-day, low-adoption portfolio analysis without tests or CI.
06 · Timeline
- Mar 10, 2017Joined GitHub
- Jan 21, 2022Created Geolocation-App-Python-Streamlit-Folium — This is an app to find Geo Location data based on the IP address of the device
- Dec 24, 2024Created Log-Realtime-Analysis — A scalable architecture for real-time log processing and visualization. Built with a Kafka-Spark ETL pipeline, DynamoDB for storing aggregate real-time metrics, and Python Dash for
- May 29, 2025Created Customer-Churn-Analysis
- May 29, 2025Most recent push to Customer-Churn-Analysis
07 · Compare
08 · Rubric
How this score was produced
Overall = Σ (category × weight) + gentle top-end curve
Tier thresholds
▸ How the pipeline works
- 01Scrape.Pull every non-fork repo pushed in the last 90 days, plus your contribution calendar, followers, and language byte counts — straight from GitHub's REST & GraphQL APIs.
- 02Triage.A small model reads every repo's file tree + README and picks the 20 files per repo that actually reveal how you code.
- 03Grade each repo. All repos run in parallel through a fast scoring model that reads the picked files and rates each one independently on Impact, Quality, and Depth — with evidence citations.
- 04Aggregate. A larger reasoning model combines the per-repo scores with server-computed stats (heatmap, commit cadence, language entropy, follower count) to produce the 6-dimension profile score + roasts.
- 05Correct.Deterministic server-side checks enforce anchor-scale floors (e.g. a profile with 2,000+ public commits can't score 30 Consistency) and recompute the final verdict.
~90 seconds per profile, ~$0.25 in compute. Total of ~240 files read across your top-12 repos. One rating per GitHub account per day.
▸ Data sources & caveats
- Heatmap & commit totals: GitHub GraphQL
contributionsCollection— covers the last 365 days, includes private repos when the user has opted in (default). - Language %: byte totals across the top 30 owned non-fork repos.
- Curve: a small upward nudge centered on raw score ≈ 70, capping at 100. Prevents specialists from being unfairly penalised for narrow breadth.
- Anchor corrections: when server-measured signals (e.g. privateWorkLikely, multiRepoVolume, follower count) mandate a minimum category score, the aggregation step enforces it. These are signal-conditional, not identity-based floors.