Skip to content
View YusefSyed's full-sized avatar

Highlights

  • Pro

Block or report YusefSyed

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
YusefSyed/README.md

Yusef Syed. Student at the University of Toronto. Building apps and developer tools. Interested in AI evaluation.

Portfolio LinkedIn Email me Resume PDF

Hey, I'm Yusef. I'm in my first year at the University of Toronto, intending Math + CS. I spend a lot of my time building apps and working on AI evaluation.

  • ⌚ I built the Apple Watch client for Providence with my team at Hack the North.
  • 🧪 My eval lab looks at what agents actually did when tool calls fail or evidence is incomplete.
  • 🌱 I'm looking for supervised research and Summer 2027 internships. Email me.

💻 Projects & experiments

agent-eval-mutation-lab: tests agent actions, partial failures, and missing results. Python. Providence: my native Apple Watch client for our Hack the North team project. Swift. tiraz-garment-completion: annotation-only PyTorch experiment with calibration and missing-context tests. agent-proof: reviewer-selected checks and redacted reports. TypeScript CLI. shiftproof: local scheduling prototype with independent constraint checks. Python. callreclaim-webmcp: a synthetic missed-call demo where the owner decides. TypeScript.

See all my repositories

🧩 Work I've contributed to

Microsoft Agent Lightning: merged shutdown race fix. Inspect Scout: merged model-usage accounting fix. Kornia: merged empty accelerator tensor fix. Microsoft PyRIT: merged package-hallucination techniques. WandB RAI Toolkit: merged default scorer category fix. PyTorch: open CPU/CUDA incomplete-gamma shape-gradient PR, not yet merged.

Each card links to my contribution. Green labels indicate merged PRs; the PyTorch PR is still open.

As of September 28, 2026, I have 12 merged upstream PRs across six projects, including Apache Arrow slicing and Agent Lightning timeout-report retries.

Contribution details and evidence

📱 My apps

The public CallReclaim Agent Desk is a separate synthetic demo with no messaging backend. ShiftProof is also a local prototype using synthetic data.

🛠️ What I work with

Languages: Python · TypeScript · JavaScript · Swift · SQL

Apps: React · React Native · Expo · Next.js · SwiftUI · watchOS

Experiments: PyTorch · NumPy · Inspect · pytest

Infrastructure: PostgreSQL · SQLite · Supabase · Docker · GitHub Actions

🔬 Research notes

The eval lab includes a frozen 624-trial local-model study. Invalid-output rates differed between conditions, so I withheld improvement claims and kept the descriptive results and missing-output analysis.

The Tiraz study uses garment annotations, not images. Its three seeded models achieved 58.5–59.0% top-1 against a 52.7% baseline on 962 held-out groups; removing context exposed a large drop in prediction-set coverage.

Methods, results, and limitations · PyTorch local validation supplement

📬 Availability

I'm available full-time May–August 2027, and open to supervised research during the school year that fits my coursework. I'm a U.S.–Canadian dual citizen, open to Canada, the U.S., relocation, and remote work. Expected graduation: May 2030.

Résumé · Email · LinkedIn

Typing intro and compact grids inspired by DenverCoder1.

Pinned Loading

  1. agent-eval-mutation-lab agent-eval-mutation-lab Public

    Reproducible experiments that check what tool-using agents actually did, including partial failures and unknown outcomes.

    Python

  2. tiraz-garment-completion tiraz-garment-completion Public

    A PyTorch garment-completion study measuring accuracy, calibration, and uncertainty when outfit context is missing.

    Python

  3. agent-proof agent-proof Public

    Run reviewer-selected checks and keep redacted, traceable evidence. A dependency-free TypeScript CLI.

    TypeScript

  4. callreclaim-webmcp callreclaim-webmcp Public

    A WebMCP missed-call desk where agents prepare replies and owners decide. Interactive demo with synthetic data.

    TypeScript

  5. shiftproof shiftproof Public

    A local volunteer-scheduling agent with independent constraint checks and coordinator approval. Synthetic-data prototype.

    Python