This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| #!/usr/bin/env python3 | |
| """ | |
| Banking77: frozen embeddings + logistic regression for 77-class intent detection. | |
| Uses the Banking77 dataset (3,080 test queries across 77 banking intents) | |
| to show what a simple pipeline can do: convert text to vectors with a | |
| pre-trained embedding model, then train an sklearn LogisticRegression on top. | |
| The embedding model is used as-is — no fine-tuning, no GPU needed. | |
| The classifier + scaler weigh about 642 KB. The embedding model |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| # This allows you to easily filter hundreds of jobs listings | |
| # to quickly determine which are the best fit for your resume | |
| # It enriches job listings descriptions by doing two things: | |
| # 1) extracts structured data from the description (eg. company name) | |
| # 2) getting answers about selection criteria (eg. is the job remote?) | |
| # It first gets job listing descriptions from a job_listings table | |
| # then assembles a prompt using a template, a resume and the | |
| # job listing |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| -- SQLite | |
| -- Sometimes the json output from GPT is not valid | |
| -- -> This query excludes invalid JSONs (about 1 in 150 for the data tested) | |
| SELECT gi.job_id, | |
| json_extract(gi.answer, '$.fit_for_resume') AS fit_for_resume, | |
| json_extract(gi.answer, '$.company_name') AS company_name, | |
| json_extract(gi.answer, '$.how_to_apply') AS how_to_apply, | |
| json_extract(gi.answer, '$.fit_justification') AS fit_justification, | |
| json_extract(gi.answer, '$.available_positions') AS available_positions, | |
| json_extract(gi.answer, '$.remote_positions') AS remote_positions, |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| # This connects to an "Ask HN: Who is hiring?" page | |
| # eg. https://news.ycombinator.com/item?id=39217310&p=1 | |
| # Gets every top-level job listing and saves it to a sqlite3 db (job_listings.db) | |
| # every entry has id, listing text, listing html and source url | |
| # (html is saved to preserve link urls, which are usually redacted within the text) | |
| # It handles pagination (follows More link at the bottom of the page) | |
| import requests | |
| from bs4 import BeautifulSoup |