Skip to content

Instantly share code, notes, and snippets.

@robertefreeman
Created August 19, 2025 20:27
Show Gist options
  • Select an option

  • Save robertefreeman/9bd89fdedfa0d27e4f87a863f3f7db02 to your computer and use it in GitHub Desktop.

Select an option

Save robertefreeman/9bd89fdedfa0d27e4f87a863f3f7db02 to your computer and use it in GitHub Desktop.
Hypernote Feedback - 8/19/2025

Hypernote Feedback (from a voice note stream-of-consciousness)

Author: @robertefreeman
Date: 2025-08-19

Note on scope and tone:

  • Everything here reflects one person’s usage, preferences, and opinions. It’s not a directive, spec, or a demand that things must change.
  • Treat these as suggestions and considerations for how I think Hypernote could work best for my workflows. Your team may have data or constraints that point in different directions.

Key themes from one user's perspective (not exhaustive)

  • Transcription as a true transcript (verbatim), with smaller chunking, and avoiding content loss when stopping recording.
  • Clear separation between “transcript” (verbatim) and “summary” (structured/smoothed), with independent autonomy/creativity controls.
  • Smaller, smarter chunking (5–10s) plus append-replace merging to reduce gaps/duplication.
  • Avoid gating open-source STT models behind Pro if they’re not hosted by you; gating cloud inference makes more sense than gating local OSS availability.
  • Templates/UI: clearer selection visibility, attachments and context scoping (e.g., resumes/JDs for interviews), larger/resizable system-instructions box.
  • Organization that prioritizes Who/What over Date; folders and tags matter; enable multi-day continuity without conflating audio.
  • Providers/models: consider adding Anthropic by default; clarify Hyper Cloud; help less-technical users with endpoints; sensible defaults.
  • Improve transcript pane formatting (line breaks, timestamps), record button affordances, and prevent summary re-gen from erasing user-asked detail.
  • Roadmap ideas: chat across multiple notes/tags with attachments, multimodal analysis, voice-to-voice with TTS, mobile apps, broader integrations (Notion, Google, OneNote).
  • Docs/website: align docs to settings, update visuals, refine Pro features copy, and cultivate a community for templates/guides.

Environment and context (for interpretation)

  • Hardware and usage:
    • Apple Silicon (e.g., M1 Pro, 32–48GB RAM). I meet on desktop, laptop, and phone. Expecting good performance headroom on modern Macs and newer mobile devices with NPUs/TPUs.
    • I favor local and self-hosted setups; I run models via LightLM AI Gateway (mix of local and API endpoints).
  • Models I tried/observed:
    • STT: Whisper small locally; I’m interested in “pro” models like Perrakey, VoxTroll, Vostrel, etc.
    • LLM: Using Meta’s Llama 4 Maverick via Meta’s API in my setup. I know this choice can color my results.
    • Hyper local model: Your “Hyper LLM” (saw references like “Hydrol/Hyper-all-in”), appears to be a 4-bit quantization (mention of “Quinn 317”), trained long enough ago to be traceable (W&B). Feels appropriate for speed/size; less so for rich cross-note analysis or multimodal.
  • Default app settings:
    • Autonomy defaulted to “autonomous.” I switched to “full autonomy” to see re-framing behavior.
  • Meeting platforms: Zoom, Teams, occasional Google Meet, plus some others historically (Cisco, BlueJeans).

Transcription: behavior, quality, and controls

1) Transcript as a true transcript (verbatim)

  • Current observation:
    • The “transcript” looks filtered/smoothed and driven by larger time gaps rather than literal output.
  • Why it matters to me:
    • Harder to audit errors and see exactly what was said if the timeline is massaged.
  • Suggestions:
    • Make the transcript window literal/verbatim by default.
    • Reserve smoothing/formatting/narrative structuring for the “summary” pane.

2) Chunking/streaming: smaller, consistent, and merge-clean

  • Observations:
    • Long gaps seem to be required before output; talking straight through yields delayed or missing chunks.
  • Suggestions:
    • Use smaller fixed chunk sizes (e.g., 5–10 seconds).
    • Apply a local aggregator to reconcile overlaps/consecutive chunks.
    • Use “append-replace” merges rather than naive append to keep the transcript clean.

3) Stopping recording should not lose in-flight content (could be a bug)

  • Current observation:
    • Hitting Stop can cut off content if the last chunk hasn’t appeared yet; I can lose 30–60+ seconds.
  • Expectation (from my perspective):
    • Stop should stop capturing more audio but allow all recorded audio to finish transcribing.
  • Suggestion:
    • Treat Stop as “end of input” while permitting queued processing to complete so nothing recorded is lost.

4) Light smoothing in transcript, configurable

  • I’m open to optional filler-word removal/light smoothing, but would prefer:
    • A toggle for aggressiveness.
    • Smaller chunks + reconciliation to preserve cadence and detail.

5) Transcription models and gating

  • I’d like to try pro STT models (Perrakey, VoxTroll, Vostrel).
  • Opinion:
    • Gating OSS models behind Pro (when not hosted by you) can feel off from a user’s perspective.
    • Gating cloud inference endpoints makes sense; gating local OSS choice feels less aligned with user expectations.

6) Transcript pane formatting

  • The transcript pane becomes a large, fatiguing block of text.
  • Suggestions:
    • Insert line breaks between sessions/chunks and when I stop/restart.
    • Add timestamps per chunk/segment to aid navigation.
    • For multi-speaker meetings, longer-term: speaker/time structuring would help.

7) Record button UI

  • I like the audio level animation for feedback.
  • The ovular control feels odd to me (subjective).
  • No strong alternative proposed—flagging the UX feel.

Summary panel and autonomy

1) Separate autonomy/creativity settings for transcript vs summary

  • It’s unclear whether the autonomy setting applies to transcript, summary, or both.
  • Suggestions:
    • Independent controls (even if in an “advanced” section), labeled clearly about what each affects.

2) Summary too high-level by default; chat help works but can be fragile

  • I used chat to request “more detail,” which worked, but re-recording/re-generation collapsed the detail back to a high-level structure.
  • Suggestions:
    • Provide a way to “lock in” desired detail or a persistent instruction so re-gens keep that density.
    • Template/system-instruction-level controls for target granularity.

3) Non-linear summary updates

  • If I add info later that logically belongs earlier, I’d love the agent to reposition it rather than just append.
  • Suggestion:
    • Allow non-linear restructuring when the system re-generates the summary so content lands in the right section.

Templates, context, and attachments

1) Clearer template selection on the main page

  • On the main interface, it isn’t obvious which template is selected for a given note.
  • Suggestions:
    • Show the selected template prominently (or via hover/label).
    • Consider context-aware defaults (e.g., Zoom vs Voice Note may imply different defaults).

2) System instructions UI size

  • The system-instruction box is ~3 lines tall; modern prompts are often lengthy.
  • Suggestion:
    • Make the area resizable and/or larger by default.

3) Template scoping and external context

  • Use case: Job Interview template.
    • I want to attach a job description and a candidate resume, plus links/docs, without polluting the core note.
  • Suggestions:
    • Allow attaching files/URLs to a note and/or to a template’s context.
    • Clarify where attachments are visible to the model: template-level, note-level, chat-only.
    • In chat, support uploads and referencing attachments with scope toggles (template only, this note, summary, etc.).
  • Rationale:
    • Interviews: role specifics + company criteria.
    • Customer work: collateral (e.g., global business plan, MedPIC/Challenger frameworks, contracts/artifacts), so I can analyze across contexts.

Organization, folders, tags, and multi-day work

1) Folders and tags feel essential

  • I saw “folders coming”—that would help a lot.
  • My workflow:
    • I prioritize “Who/What” first, “When” second.
    • I often return later to add thoughts to a prior meeting.
  • Suggestions:
    • Organize by Who/What (company/project/people) then Date.
    • Let me add a new session tied to a previous meeting without conflating the original audio timeline.
    • Tags across notes (e.g., “Caterpillar” + opportunity number) and the ability to chat across all items with a given tag/folder.

Intelligence settings, providers, and defaults

1) Providers

  • Defaults: OpenAI, Google Gemini (Flash/FlashLite), OpenRouter.
  • Suggestion:
    • Consider adding Anthropic as a first-class option (Haiku can be cost-effective here).
    • For “Other” endpoints, add guardrails (auto-add v1, handle chat/completions paths) to reduce setup errors for less-technical users.

2) Hyper LLM (local) and Hyper Cloud

  • Local “Hyper LLM” feels right for responsiveness; perhaps limited for deeper/multimodal tasks.
  • Curiosity/questions (not demands):
    • Is Hyper Cloud just a bigger/faster variant of the local model, or different models entirely?
    • Will there be quotas/usage caps, and multi-model options?
    • How do you plan to handle cost variability across users?
  • Suggestion:
    • A clear reference page/table explaining model choices, strengths, caps, and pricing interactions.

3) Model choice impact and defaults

  • My high-level summaries may be influenced by my chosen model (Llama 4 Maverick).
  • Suggestion:
    • Provide defaults that tend to produce balanced detail for note-taking, and allow per-template model/behavior settings.

Multimodal, voice-to-voice, and richer interactions

  • Multimodal:
    • I’d like to attach slides/images and ask for summaries/analyses.
    • Data viz like “chart X from these notes” would be great.
  • Voice-to-voice chat:
    • With capable STT (e.g., Vostrel Small) and a lightweight TTS (Kokoro, Chatterbox), I could talk to my notes conversationally.
    • Longer-term, on-device voice chat feels feasible on modern hardware.
  • Chat across folders/tags:
    • “Talk to everything tagged Caterpillar,” including attached collateral.

Integrations and platforms

  • Obsidian integration fits a privacy-focused audience.
  • Additional integrations I’d personally find valuable:
    • Notion, Google Docs/Drive, Gmail, OneNote, etc.
  • Mobile:
    • iOS/Android apps would help for meeting capture and quick notes; desktop-first is okay initially, but cross-device is important.

Source Analysis and “Ask AI” UX

  • “Source analysis” appears to be behind a compute decision; demo/screencap may be outdated.
  • If “Ask AI” is coming:
    • Align visuals and flows (mental model similar to Google Docs AI helpers).
  • Suggestion:
    • Keep feature names/status and tutorials in sync to avoid confusion.

Pricing, Pro features, and gating strategy

  • Opinionated take:
    • Gating cloud inference behind Pro makes sense.
    • Gating locally-run open-source STT models behind Pro feels misaligned (as a user paying to use my own hardware/OSS).
  • Pro page copy:
    • The “Other small things” bullet undersells value.
    • Emphasize: “We include an AI inference runtime (local + cloud options)” if that’s accurate.
    • Trials/credits could help (I understand cost variability makes this non-trivial).
  • Community access (e.g., Discord):
    • I’d keep this open; cultivating champions/templates/integrations/blogs can compound value.

Documentation and website feedback

1) Docs alignment and discoverability

  • I’d love docs that mirror Settings and key workflows 1:1.
  • Explicit pages for:
    • What autonomy affects (transcript vs summary vs both).
    • What re-generation does to user-edited summaries.

2) Visual theming and logo notes (subjective)

  • The site uses warm peach/cream tones; some screenshots/assets look cold (white/blue frames), reading like dropped-in assets rather than themed.
  • Logo kerning/spacing:
    • The lightning bolt between R and N is cool; spacing elsewhere feels looser by contrast.
    • I’m not a designer; just a subjective note.

Specific UI/UX suggestions (subjective)

  • Transcript pane:
    • Line breaks between sessions/chunks.
    • Timestamps on segments.
  • Summary:
    • Preserve user-requested detail across re-gens; allow “lock” or persistent directive.
  • Template selection:
    • Show the selected template on the main list.
    • Consider context-aware defaults.
  • Attachments:
    • Allow attachments (files/links) with scope controls; let chat reference them.
  • System instructions:
    • Resizable/larger prompt box; support long-form instructions.
  • Autonomy:
    • Separate controls with clear scope tooltips.
  • Stop-recording behavior:
    • Finish processing recorded audio before finalizing.

Models, STT/TTS, and performance notes

  • STT:
    • Whisper small is usually fast on Apple Silicon; my hiccups likely relate to execution/timing.
    • I’d like to try Perrakey, VoxTroll, Vostrel without Pro gating if they’re local/OSS.
  • TTS:
    • Pluggable TTS (Kokoro, Chatterbox) could enable voice responses/conversation.
  • Local model capability:
    • Your local model feels like a sensible speed/size compromise. For bigger tasks (cross-note, multimodal), users may prefer occasional cloud calls.
  • Device headroom:
    • Even older M1 Pro machines are capable; newer M3/M4 and modern Androids increase headroom.

Integrations and ecosystem: community as a force multiplier

  • Community strategy suggestions:
    • Encourage users to build/share templates, write guides, and showcase integrations.
    • Highlight community contributions in blogs; consider a templates exchange (not necessarily paid).
  • Education:
    • Many users are endpoint-/API-setup shy; first-class provider buttons and guardrails help a lot.

Bugs, enhancements, and open questions (framed as suggestions)

Possible bug (from my perspective)

  • Stopping recording appears to truncate transcript if the last chunk hasn’t emitted yet.
    • Suggestion: Let queued processing finish so all recorded audio is transcribed.

Possible enhancements

  • Verbatim transcript with optional light smoothing and controls.
  • Smaller chunking (5–10s) with append-replace merging and overlap reconciliation.
  • Separate autonomy settings for transcript and summary; clarify scope in the UI.
  • Preserve user-directed detail across re-gens; allow “lock”/persistent directives.
  • Non-linear summary restructuring (place late info in the right section).
  • Template selection visibility; context-aware defaults.
  • Resizable system-instructions box; attachments in notes/templates; chat access with scope controls.
  • Transcript pane: line breaks, timestamps.
  • Add Anthropic as a first-class provider; smooth “Other” endpoint setup.
  • Explain Hyper Cloud models, limits, and pricing approach.
  • Add mobile apps; broaden integrations (Notion, Google, OneNote, Gmail).
  • Multimodal analysis; voice-to-voice chat with TTS.
  • Update “Source analysis” visuals; clarify “Ask AI” flows.
  • Refine Pro page; reconsider gating OSS STT locally; consider trials.

Open questions (curiosity)

  • Does the autonomy slider affect transcript, summary, or both?
  • Can user-set detail preferences persist across re-gens?
  • Will Hyper Cloud offer multiple model choices and quotas?
  • Roadmap for folders/tags and cross-note chat?
  • Plans for multimodal and voice-to-voice as first-class features?

Possible near-term priorities (subjective)

  1. Address stop-recording truncation so recorded audio always finishes transcription.
  2. Make the transcript verbatim and adopt smaller-chunk pipeline with clean merging.
  3. Separate transcript vs summary behavior and autonomy controls; clarify what each affects.
  4. Preserve user-requested summary detail across re-gens; allow “lock” or persistent directives.
  5. Show selected template; enlarge/resizable system-instruction editor; enable attachments + scoping.
  6. Consider ungating local OSS STT models from Pro while keeping Pro gating for hosted/cloud inference.
  7. Add timestamps/line breaks in the transcript pane.

Closing

Overall, I like the product direction and design language, and I appreciate the privacy-first/local inference stance. My biggest personal pain points are transcript fidelity (and not losing content), clarity/control over autonomy and re-generation behavior, and OSS model gating. The areas I’m most excited about are cross-note/tag chat with attachments, multimodal, and voice-to-voice—features that could make Hypernote a central productivity hub.

Please take all of this as one user’s perspective to inform your roadmap—not as a directive.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment