frisket
githubcontact

Frisket is an open-source AI spreadsheet for reporting! for research! for investigations!

You can drop in documents, audio, video, emails, spreadsheets, etc. Ask questions, get structured data back.

Imagine the following:

  • Excel with (actually good) AI integration
  • OpenRefine with magic powers
  • Pinpoint without being tied to Google

Love privacy? You can run Frisket on your machine, only using private local models and never touching the outside world.

Love corporations? Frisket also has plenty of pathways to feed your favorite providers like OpenAI, Anthropic, Google and OpenRouter.

Using Frisket

Non-Python people can sign up for the hosted version.

Python people can install and run Frisket with the command below. For more details about hosting on the web, Docker, sidecars, GPUs, extra features, etc read up on GitHub.

pip install 'frisket-data[standard]'
frisket ./my-workspace

Why I made Frisket

Modern-day investigative journalists and other researchers have to tackle large caches of documents to uncover stories and secrets all the time. When you need to wrangle ten thousand court records, transcribe twenty million hours of meeting, or parse out nine thousand TikTok videos, you could read every word and watch every second, but AI can do a lot to speed things up.

I teach folks how to do this, but you're coaxing best-effort solutions out of a hundred and one different (wonderful) tools. If only there was a beautiful, perfect answer!

Earlier this summer I heard about a super-interesting NYT tool and I thought: okay, well, time to make one for the rest of us.

And now we have Frisket.

What Frisket does

While no one has used Frisket for much of anything yet (that's your job!), here are the kinds of great stories you could use Frisket for:

Just look at Andrew Deck's writeup on Pulitzer Prine winning stories and pretend I'm behind you repeating "okay yes, Frisket can do that" at the end of each paragraph.

How do you use Frisket?

It looks like Excel or Google Sheets, and works similarly.

You upload your content, you click the tools you want (extract, transcribe, detect topic changes, etc), and then you bask in the results. It's definitely overwhelming, so there's a Guide button that walks you thought eight! different workflows as gentle tutorials.

In the example below, I uploaded some PDFs that may or may not have been court documents. I then clicked the "extract" button, said "give me the plaintiff, the defendant, the date and a summary," and it did – even though I couldn't spell defendant correctly.

Who are you?

I'm Jonathan Soma, Knight Chair in Data Journalism at Columbia's Journalism School, where I direct two data journalism programs – one is twelve months long and the other is ten weeks. I write software for journalism and lecture a lot about how to use AI and evaluate AI in the journalism industry.

If you love Frisket you can either come to the J-School or donate to my cat rescue. If you hate Frisket you should donate to the rescue so I have more time to tell Codex to fix bugs.

You can find me via email at js4571@columbia.edu or on Twitter at @dangerscarf.