Skip to content

Commit 442d087

Browse files
committed
Enhance company configuration and scripts for improved user filtering
- Added a new script command "lookup" to package.json for candidate lookup functionality. - Updated companyConfig in company.ts to refine tier descriptions and adjust weights for seniority_fit and hireability criteria. - Implemented NYC-only filtering in review-batch.ts and best_rated_to_txt.ts to focus on local candidates. - Improved LinkedIn profile fetching logic in linkedin-research.ts to streamline search queries and added guidelines for query construction.
1 parent f850c4a commit 442d087

9 files changed

Lines changed: 457 additions & 20 deletions

File tree

‎outreach-agent-guide.md‎

Lines changed: 101 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,101 @@
1+
# Outreach Agent Guide
2+
3+
## Overview
4+
5+
Help Moritz reach out to engineering candidates found by the GitHub scraper. For each candidate: fetch their data, draft a personalized connection message, track in the Google Sheet, and send via LinkedIn/X/email.
6+
7+
## 1. Fetch Candidate Data
8+
9+
```bash
10+
npm run lookup <github-username-or-linkedin-slug>
11+
```
12+
13+
Returns JSON with: name, email, company, bio, location, rating, archetype, links (LinkedIn, X, blog), linkedinSummary, currentCompanyInsights, webResearch, criteriaScores, criteriaReasonings, and graph discovery data (discoveredVia, parentRatings, depth).
14+
15+
Key fields for message drafting:
16+
17+
- `linkedinSummary` - full career history, companies, roles, dates
18+
- `criteriaReasonings` - per-criterion analysis (startup_experience, ai_agent_experience, builder_signal, etc.)
19+
- `webResearch` - web research summary about the person
20+
- `currentCompanyInsights` - company headcount, growth trends (useful for hireability context)
21+
- `bio` / `xBio` - GitHub and X bios
22+
- `blog` - personal website
23+
- `topReferrer` - the highest-rated person who led us to this candidate in the GitHub graph, with their full name already resolved. Use this to mention a shared connection, e.g. "I noticed you follow Ishaan Dey on GitHub - small world!"
24+
- `discoveredVia` - "following" (parent follows them) or "followers" (they follow the parent)
25+
- `parentRatings` - all parents in the graph (topReferrer is the best one)
26+
27+
## 2. Web Search
28+
29+
Before drafting, do a quick web search for the person to find recent news, blog posts, or projects not captured in the DB. This helps personalize the message.
30+
31+
## 3. Draft Connection Message
32+
33+
**LinkedIn connection requests have a 300 character limit.** Always verify the exact character count programmatically (`echo -n "message" | wc -c`) - never estimate. LLMs are bad at counting characters.
34+
35+
Guidelines:
36+
37+
- Reference something specific and real about their work - check their actual website/repos, don't make claims you can't verify
38+
- Mention Rogo is backed by Sequoia and Thrive Capital
39+
- Keep it concise and to the point - no fluff or buzzwordsused th
40+
- End with a soft ask to chat with Gabe (Rogo CEO)
41+
- Never use emojis
42+
- Fetch existing outreach messages from the Google Sheet (column B) and match the tone and style. Do NOT use hardcoded examples - always pull real ones from the sheet.
43+
44+
Personalization ideas (use what's relevant):
45+
46+
- A specific project/product they built with a concrete metric (e.g. "40k+ installs")
47+
- Their founding/startup experience ("many ex-founders like yourself here")
48+
- Number of mutual LinkedIn connections ("I noticed we have X mutuals")
49+
- How they were discovered via the graph ("I noticed you follow [person] on GitHub")
50+
- Their career background if relevant to Rogo (finance, AI, productivity tools)
51+
52+
## 4. Google Sheet Tracking
53+
54+
**Sheet ID**: `1HJeeQiF0KBf-S5PNLHMJhtdVbN6qoPjG6qUtvFuwgrA`
55+
**Sheet name**: `Tabellenblatt1`
56+
**Auth**: `gcloud auth print-access-token`
57+
58+
Current columns (A-O):
59+
60+
- A: Name
61+
- B: Outreach Messages (the message sent to this person)
62+
- C: Current/past companies
63+
- D: Why interesting / potential role
64+
- E: Scraper Score
65+
- F: LinkedIn DM Moritz (date)
66+
- G: X DM Moritz (date)
67+
- H: Email Moritz (date)
68+
- I: Status
69+
- J: Location
70+
- K: Email Address
71+
- L: LinkedIn
72+
- M: GitHub
73+
- N: Twitter
74+
- O: Other links
75+
76+
### Read sheet
77+
78+
```bash
79+
curl -s "https://sheets.googleapis.com/v4/spreadsheets/1HJeeQiF0KBf-S5PNLHMJhtdVbN6qoPjG6qUtvFuwgrA/values/Tabellenblatt1" \
80+
-H "Authorization: Bearer $(gcloud auth print-access-token)"
81+
```
82+
83+
### Append a row
84+
85+
```bash
86+
curl -s -X POST \
87+
"https://sheets.googleapis.com/v4/spreadsheets/1HJeeQiF0KBf-S5PNLHMJhtdVbN6qoPjG6qUtvFuwgrA/values/Tabellenblatt1:append?valueInputOption=USER_ENTERED" \
88+
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
89+
-H "Content-Type: application/json" \
90+
-d '{"values": [["Name", "outreach msg", "Companies", "Why interesting", "Score", "Feb 13", "", "", "reached out", "Location", "[email protected]", "https://linkedin.com/in/slug", "https://github.com/user", "@handle", ""]]}'
91+
```
92+
93+
Always refetch the sheet before adding rows to check current layout and avoid duplicates.
94+
95+
## 5. Email Outreach
96+
97+
Email can be sent from [email protected] via MCP (setup in progress - check if the email MCP server is available before attempting).
98+
99+
## 6. Company Context
100+
101+
Rogo is a Series C AI startup backed by Sequoia and Thrive Capital. We build productivity software for investment banking and private equity: AI-powered presentation generation, Excel automation, research agents, and financial data tools. CEO is Gabe Stengel (@GabeStengel). Based in NYC.

‎package.json‎

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -48,7 +48,8 @@
4848
"links": "tsx src/graph-scraper/scripts/print-links.ts",
4949
"review": "tsx src/graph-scraper/scripts/review-batch.ts",
5050
"scrape-one": "tsx src/graph-scraper/scripts/scrape-one.ts",
51-
"re-rate": "tsx src/graph-scraper/scripts/re-rate-users.ts"
51+
"re-rate": "tsx src/graph-scraper/scripts/re-rate-users.ts",
52+
"lookup": "tsx src/graph-scraper/scripts/lookup-candidate.ts"
5253
},
5354
"keywords": [],
5455
"author": "",

‎src/config/company.ts‎

Lines changed: 13 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -48,7 +48,7 @@ export const companyConfig: {
4848
label: 'Startup Experience',
4949
weight: 3,
5050
tiers: {
51-
0: 'No startup experience or non-technical startup roles',
51+
0: 'No startup experience or non-technical startup roles. Also use for someone who has spent their entire career (10+ years) at one or two large corporations (e.g., Bloomberg, IBM, Oracle, Microsoft, Google) without any startup involvement - these candidates are very unlikely to thrive in a fast-moving Series B environment.',
5252
1: 'Worked at a startup in a hands-on engineering role, but not a well-known or fast-growing one. Indie hackers and solo SaaS builders without significant traction also fall here.',
5353
2: 'Founding engineer or early engineer at a startup with some validation (known investors, meaningful revenue, or growing team)',
5454
3: 'Founded or co-founded a productivity/AI/fintech startup with strong validation (tier-1 VC funding, acquisition, significant traction)',
@@ -61,8 +61,8 @@ export const companyConfig: {
6161
tiers: {
6262
0: 'No AI or agent experience',
6363
1: 'General interest, courses, or minor AI or agent projects',
64-
2: 'Built AI-powered tools or applied AI or agent in a real product',
65-
3: 'Shipped AI agents, RAG systems, text-to-SQL, document processing, or research automation in production',
64+
2: 'Built AI-powered tools or applied AI or agent in a real product. Corporate ML/inference infrastructure (model serving, ML pipelines) at large companies falls here, not tier 3.',
65+
3: 'Shipped AI agents, RAG systems, text-to-SQL, document processing, or research automation in production. Must be building the AI-powered product itself, not just the infra/platform underneath it.',
6666
},
6767
},
6868
{
@@ -134,7 +134,7 @@ export const companyConfig: {
134134
{
135135
key: 'seniority_fit',
136136
label: 'Seniority Fit',
137-
weight: 1,
137+
weight: 2,
138138
tiers: {
139139
0: 'VP/C-suite at a well-known or large company, famous tech leader, tenured professor - way too senior for a Series B startup',
140140
1: 'Director at a large company, engineering manager whose recent roles are primarily people management. Staff/tech lead at a big company with managerial responsibilities also falls here - they may struggle to go back to pure IC work.',
@@ -156,18 +156,18 @@ export const companyConfig: {
156156
{
157157
key: 'hireability',
158158
label: 'Hireability',
159-
weight: 4,
159+
weight: 6,
160160
tiers: {
161161
0: 'CEO/CTO/co-founder/VP at a company that is clearly growing (positive headcount growth, >10 employees, or raised significant funding recently). Use company insights data if available. Also: anyone in a senior position (Principal, Staff, Distinguished, Director+) at a rocket-ship AI company (Anthropic, OpenAI, Thinking Machines, Cursor, etc.) - these people are extremely well-compensated and will not leave. These people will not leave their company.',
162-
1: 'Co-founder/exec at a funded startup with moderate or unknown growth, or C-suite at an established company. Also: junior/mid-level IC engineer at a rocket-ship AI company (Anthropic, OpenAI, Cursor, etc.) where leaving would be irrational. Also: someone who just started a new role or company (<6 months ago) - they are in the honeymoon phase and very unlikely to leave.',
162+
1: 'Co-founder/exec at a funded startup with moderate or unknown growth, or C-suite at an established company. Also: junior/mid-level IC engineer at a rocket-ship AI company (Anthropic, OpenAI, Cursor, etc.) where leaving would be irrational. Also: someone who just started a new role or company (<6 months ago) - they are in the honeymoon phase and very unlikely to leave. Also: someone who has been at the same large company for 10+ years - they are deeply embedded and very unlikely to leave for a startup.',
163163
2: 'Founder of a small/stagnating/early-stage company (<5 employees, no/negative growth in company insights), recently exited founder, someone whose company shut down. Also use for someone stuck at a tiny company (1-3 employees, no growth) for 3+ years - this signals they may be comfortable/complacent rather than ambitious. Serial indie hackers/bootstrappers who have been running their own small projects for 5+ years are also unlikely to join a venture-backed startup.',
164164
3: 'Employee (not founder/exec), IC engineer at a normal company, or someone clearly between roles and open to new opportunities. Not at a rocket-ship company. Not stuck at a stagnant company for years. Not a serial indie hacker.',
165165
},
166166
},
167167
{
168168
key: 'role_fit',
169169
label: 'Role Fit',
170-
weight: 1,
170+
weight: 2,
171171
tiers: {
172172
0: 'Not a relevant engineering role (PM, designer, researcher only, or no engineering background). Also: robotics, embedded systems, hardware, computer vision, or other non-web engineering.',
173173
1: 'Adjacent engineering role (data engineer, DevOps, ML researcher, mobile-only, or primarily ML/CV engineer who does some web work on the side)',
@@ -267,6 +267,12 @@ export const companyConfig: {
267267
'https://github.com/kamath',
268268
'https://github.com/adamcohenhillel',
269269
'https://github.com/tommoor',
270+
'https://github.com/juliusmarminge',
271+
'https://github.com/mfts',
272+
'https://github.com/raghavpillai',
273+
// other ppl i respect
274+
'https://github.com/samuelstroschein',
275+
'https://github.com/mitsuhiko',
270276
],
271277

272278
// The full LLM rating prompt (static part).

‎src/graph-scraper/core/scraper-helpers/linkedin-research.ts‎

Lines changed: 2 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -878,9 +878,7 @@ export async function fetchLinkedInProfileUsingBrave(
878878
const performFetchTask = async () => {
879879
const searchQuery = optimizedQuery
880880
? `site:linkedin.com/in/ ${optimizedQuery}`
881-
: `site:linkedin.com/in/ ${user.name || user.login} ${
882-
user.email ? `email:${user.email}` : ""
883-
} ${user.xBio || user.bio || ""} (Software Engineer)`;
881+
: `site:linkedin.com/in/ ${user.name || user.login} Software Engineer`;
884882

885883
try {
886884
if (!process.env.BRAVE_API_KEY) {
@@ -1033,6 +1031,7 @@ RULES FOR QUERIES:
10331031
5. Extract company names from bio, X bio, or company field (e.g. "@anysphere" -> "Cursor", "@vercel" -> "Vercel")
10341032
6. First name + specific company is OK if the combination is unique (e.g. "Jason Cursor Skiff CTO")
10351033
7. Only fall back to "Software Engineer" if there is truly no company info available
1034+
8. NEVER include open-source project names, library names, or GitHub repo names in the query (e.g. "htmx", "tRPC", "prisma"). Nobody puts these on their LinkedIn profile. Use the company they work at or "Software Engineer" instead.
10361035
10371036
Format your response exactly as:
10381037
REASONING: [Your detective work here]

‎src/graph-scraper/output-gen/best_rated_to_txt.ts‎

Lines changed: 7 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -17,6 +17,8 @@ const excludedArchetypes = [
1717
"None",
1818
];
1919

20+
const NYC_ONLY = process.argv.includes("--nyc");
21+
2022
interface RatedUser {
2123
_id: string;
2224
rating: number;
@@ -63,7 +65,11 @@ async function exportBestRatedToTxt() {
6365
.sort({ rating: -1 })
6466
.toArray();
6567

66-
const slicedRatedUsers = ratedUsers.slice(startIndex, endIndex);
68+
const filteredUsers = NYC_ONLY
69+
? ratedUsers.filter((u) => u.criteriaScores?.location === 3)
70+
: ratedUsers;
71+
72+
const slicedRatedUsers = filteredUsers.slice(startIndex, endIndex);
6773

6874
// Create output directory if it doesn't exist
6975
const outputDir = path.join(process.cwd(), "output");
Lines changed: 143 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,143 @@
1+
import dotenv from "dotenv";
2+
import { MongoClient } from "mongodb";
3+
dotenv.config();
4+
5+
async function main() {
6+
const client = new MongoClient(process.env.MONGODB_URI!);
7+
await client.connect();
8+
const db = client.db(process.env.MONGODB_DB);
9+
const users = db.collection("users");
10+
11+
// 1. Does "following" vs "followers" discovery direction predict score?
12+
console.log("=== SCORE BY DISCOVERY DIRECTION ===");
13+
for (const via of ["following", "followers"]) {
14+
const results = await users.aggregate([
15+
{ $match: { status: "processed", rating: { $exists: true }, discoveredVia: via } },
16+
{ $group: { _id: null, avgScore: { $avg: "$rating" }, count: { $sum: 1 }, scored40plus: { $sum: { $cond: [{ $gte: ["$rating", 40] }, 1, 0] } } } }
17+
]).toArray();
18+
if (results.length) {
19+
const r = results[0];
20+
console.log(` ${via}: avg=${r.avgScore.toFixed(1)}, count=${r.count}, 40+=${r.scored40plus} (${((r.scored40plus/r.count)*100).toFixed(1)}%)`);
21+
}
22+
}
23+
24+
// 2. Does depth predict score?
25+
console.log("\n=== SCORE BY DEPTH ===");
26+
const depthResults = await users.aggregate([
27+
{ $match: { status: "processed", rating: { $exists: true }, depth: { $exists: true } } },
28+
{ $group: { _id: "$depth", avgScore: { $avg: "$rating" }, count: { $sum: 1 }, scored40plus: { $sum: { $cond: [{ $gte: ["$rating", 40] }, 1, 0] } } } },
29+
{ $sort: { _id: 1 } }
30+
]).toArray();
31+
for (const r of depthResults) {
32+
console.log(` depth=${r._id}: avg=${r.avgScore.toFixed(1)}, count=${r.count}, 40+=${r.scored40plus} (${((r.scored40plus/r.count)*100).toFixed(1)}%)`);
33+
}
34+
35+
// 3. Does max parent rating predict score?
36+
console.log("\n=== SCORE BY MAX PARENT RATING ===");
37+
const processed = await users.find({
38+
status: "processed",
39+
rating: { $exists: true },
40+
parentRatings: { $exists: true, $ne: [] }
41+
}).project({ rating: 1, parentRatings: 1, depth: 1, discoveredVia: 1 }).toArray();
42+
43+
const buckets: Record<string, { total: number; good: number; sumScore: number }> = {
44+
"parent 55+": { total: 0, good: 0, sumScore: 0 },
45+
"parent 45-55": { total: 0, good: 0, sumScore: 0 },
46+
"parent 35-45": { total: 0, good: 0, sumScore: 0 },
47+
"parent 25-35": { total: 0, good: 0, sumScore: 0 },
48+
"parent <25": { total: 0, good: 0, sumScore: 0 },
49+
};
50+
51+
for (const u of processed) {
52+
const parentScores = (u.parentRatings || []).map((p: any) => typeof p === "number" ? p : p.rating).filter((r: any) => typeof r === "number");
53+
if (!parentScores.length) continue;
54+
const maxParent = Math.max(...parentScores);
55+
56+
let bucket: string;
57+
if (maxParent >= 55) bucket = "parent 55+";
58+
else if (maxParent >= 45) bucket = "parent 45-55";
59+
else if (maxParent >= 35) bucket = "parent 35-45";
60+
else if (maxParent >= 25) bucket = "parent 25-35";
61+
else bucket = "parent <25";
62+
63+
buckets[bucket].total++;
64+
buckets[bucket].sumScore += u.rating;
65+
if (u.rating >= 40) buckets[bucket].good++;
66+
}
67+
68+
for (const [label, b] of Object.entries(buckets)) {
69+
if (b.total > 0) {
70+
console.log(` ${label}: avg=${(b.sumScore/b.total).toFixed(1)}, count=${b.total}, 40+=${b.good} (${((b.good/b.total)*100).toFixed(1)}%)`);
71+
}
72+
}
73+
74+
// 4. Does number of parents (discovered by multiple high-scoring people) predict score?
75+
console.log("\n=== SCORE BY NUMBER OF PARENTS ===");
76+
const parentCountBuckets: Record<string, { total: number; good: number; sumScore: number }> = {
77+
"1 parent": { total: 0, good: 0, sumScore: 0 },
78+
"2-3 parents": { total: 0, good: 0, sumScore: 0 },
79+
"4-6 parents": { total: 0, good: 0, sumScore: 0 },
80+
"7+ parents": { total: 0, good: 0, sumScore: 0 },
81+
};
82+
83+
for (const u of processed) {
84+
const nParents = (u.parentRatings || []).length;
85+
if (!nParents) continue;
86+
87+
let bucket: string;
88+
if (nParents >= 7) bucket = "7+ parents";
89+
else if (nParents >= 4) bucket = "4-6 parents";
90+
else if (nParents >= 2) bucket = "2-3 parents";
91+
else bucket = "1 parent";
92+
93+
parentCountBuckets[bucket].total++;
94+
parentCountBuckets[bucket].sumScore += u.rating;
95+
if (u.rating >= 40) parentCountBuckets[bucket].good++;
96+
}
97+
98+
for (const [label, b] of Object.entries(parentCountBuckets)) {
99+
if (b.total > 0) {
100+
console.log(` ${label}: avg=${(b.sumScore/b.total).toFixed(1)}, count=${b.total}, 40+=${b.good} (${((b.good/b.total)*100).toFixed(1)}%)`);
101+
}
102+
}
103+
104+
// 5. Does "following" from a high-scorer beat "followers" from a high-scorer?
105+
console.log("\n=== COMBINED: DIRECTION + MAX PARENT SCORE ===");
106+
const comboBuckets: Record<string, { total: number; good: number; sumScore: number }> = {};
107+
108+
for (const u of processed) {
109+
const parentScores = (u.parentRatings || []).map((p: any) => typeof p === "number" ? p : p.rating).filter((r: any) => typeof r === "number");
110+
if (!parentScores.length) continue;
111+
const maxParent = Math.max(...parentScores);
112+
const via = u.discoveredVia || "unknown";
113+
114+
let parentBucket: string;
115+
if (maxParent >= 45) parentBucket = "45+";
116+
else if (maxParent >= 35) parentBucket = "35-45";
117+
else parentBucket = "<35";
118+
119+
const key = `${via} + parent ${parentBucket}`;
120+
if (!comboBuckets[key]) comboBuckets[key] = { total: 0, good: 0, sumScore: 0 };
121+
comboBuckets[key].total++;
122+
comboBuckets[key].sumScore += u.rating;
123+
if (u.rating >= 40) comboBuckets[key].good++;
124+
}
125+
126+
for (const [label, b] of Object.entries(comboBuckets).sort((a, b) => (b[1].good/b[1].total) - (a[1].good/a[1].total))) {
127+
if (b.total > 10) {
128+
console.log(` ${label}: avg=${(b.sumScore/b.total).toFixed(1)}, count=${b.total}, 40+=${b.good} (${((b.good/b.total)*100).toFixed(1)}%)`);
129+
}
130+
}
131+
132+
// 6. GitHub metadata available before scraping: do we have followers count, public repos, etc?
133+
console.log("\n=== GITHUB METADATA CORRELATION ===");
134+
// Check if we have any pre-scrape metadata
135+
const sampleWithMeta = await users.findOne(
136+
{ status: "processed", rating: { $exists: true } },
137+
{ projection: { _id: 1, followers: 1, following: 1, public_repos: 1, githubFollowers: 1, githubData: 1 } }
138+
);
139+
console.log(" Sample fields:", JSON.stringify(Object.keys(sampleWithMeta || {})));
140+
141+
await client.close();
142+
}
143+
main().catch(console.error);

0 commit comments

Comments
 (0)