all work

Muse Studio

(AI Writing Tool)

An AI writing tool grounded in a team's own docs, so drafts sound like the brand, not a chatbot.

  • 3× faster first drafts
  • 9k monthly users
  • −60% token cost

2025 · Next.js · TypeScript · pgvector · LLM API · Vercel

musestudio.app

Problem

(what hurt)

Writers were pasting brand guidelines into generic chat tools and still rewriting everything. Outputs were off-voice and impossible to fact-check.

Project goal: Content teams wanted AI help that sounds like their brand. The goal was a writing tool grounded in their own documents and style guide, with sources they can check.

Architecture

(how it fits together)

Documents are chunked and embedded into pgvector. Each request retrieves the closest passages, builds a prompt with the style guide, and streams the answer back to the editor, citing the passages it used.

Tech decisions: Retrieval over the team's documents with pgvector, responses streamed token by token from an LLM API through a Next.js route handler, and prompt caching to keep costs predictable.

  • Next.js
  • TypeScript
  • pgvector
  • LLM API
  • Vercel
(architecture)EditorStream APIpgvectorCache
app/api/draft/route.ts
export async function POST(req: Request) {
  const { prompt, docId } = await req.json();
  const context = await vectorSearch(docId, prompt, { k: 6 });

  const stream = await llm.stream({
    system: STYLE_GUIDE,
    messages: [{ role: "user", content: withContext(prompt, context) }],
  });

  return new Response(stream.toReadableStream(), {
    headers: { "Content-Type": "text/event-stream" },
  });
}

Features

(what it does)

  • Grounded answers

    Every paragraph links back to the source passage, so writers can check claims in one click.

  • Streaming editor

    Text arrives word by word inside the editor, and writers can stop or steer mid-draft.

  • Voice presets

    Teams save tone presets per channel, from support replies to launch posts.

(in use)
ship it ✓
musestudio.app/settings

Challenges

(what was hard)

  1. 01Cost control

    Prompt caching for the style guide and smarter chunking cut token spend by 60%.

  2. 02Latency

    Retrieval and generation start in parallel where possible, so the first words appear in under a second.

Results

  • 3× faster first drafts
  • 9k monthly users
  • −60% token cost

Teams write first drafts three times faster, and Muse grew to 9k monthly users without paid acquisition.

design & code by Shreya Pathak