Author: @robertefreeman
Date: 2025-08-19
Note on scope and tone:
- Everything here reflects one person’s usage, preferences, and opinions. It’s not a directive, spec, or a demand that things must change.
- Treat these as suggestions and considerations for how I think Hypernote could work best for my workflows. Your team may have data or constraints that point in different directions.
- Transcription as a true transcript (verbatim), with smaller chunking, and avoiding content loss when stopping recording.
- Clear separation between “transcript” (verbatim) and “summary” (structured/smoothed), with independent autonomy/creativity controls.
- Smaller, smarter chunking (5–10s) plus append-replace merging to reduce gaps/duplication.
- Avoid gating open-source STT models behind Pro if they’re not hosted by you; gating cloud inference makes more sense than gating local OSS availability.
- Templates/UI: clearer selection visibility, attachments and context scoping (e.g., resumes/JDs for interviews), larger/resizable system-instructions box.
- Organization that prioritizes Who/What over Date; folders and tags matter; enable multi-day continuity without conflating audio.
- Providers/models: consider adding Anthropic by default; clarify Hyper Cloud; help less-technical users with endpoints; sensible defaults.
- Improve transcript pane formatting (line breaks, timestamps), record button affordances, and prevent summary re-gen from erasing user-asked detail.
- Roadmap ideas: chat across multiple notes/tags with attachments, multimodal analysis, voice-to-voice with TTS, mobile apps, broader integrations (Notion, Google, OneNote).
- Docs/website: align docs to settings, update visuals, refine Pro features copy, and cultivate a community for templates/guides.
- Hardware and usage:
- Apple Silicon (e.g., M1 Pro, 32–48GB RAM). I meet on desktop, laptop, and phone. Expecting good performance headroom on modern Macs and newer mobile devices with NPUs/TPUs.
- I favor local and self-hosted setups; I run models via LightLM AI Gateway (mix of local and API endpoints).
- Models I tried/observed:
- STT: Whisper small locally; I’m interested in “pro” models like Perrakey, VoxTroll, Vostrel, etc.
- LLM: Using Meta’s Llama 4 Maverick via Meta’s API in my setup. I know this choice can color my results.
- Hyper local model: Your “Hyper LLM” (saw references like “Hydrol/Hyper-all-in”), appears to be a 4-bit quantization (mention of “Quinn 317”), trained long enough ago to be traceable (W&B). Feels appropriate for speed/size; less so for rich cross-note analysis or multimodal.
- Default app settings:
- Autonomy defaulted to “autonomous.” I switched to “full autonomy” to see re-framing behavior.
- Meeting platforms: Zoom, Teams, occasional Google Meet, plus some others historically (Cisco, BlueJeans).
- Current observation:
- The “transcript” looks filtered/smoothed and driven by larger time gaps rather than literal output.
- Why it matters to me:
- Harder to audit errors and see exactly what was said if the timeline is massaged.
- Suggestions:
- Make the transcript window literal/verbatim by default.
- Reserve smoothing/formatting/narrative structuring for the “summary” pane.
- Observations:
- Long gaps seem to be required before output; talking straight through yields delayed or missing chunks.
- Suggestions:
- Use smaller fixed chunk sizes (e.g., 5–10 seconds).
- Apply a local aggregator to reconcile overlaps/consecutive chunks.
- Use “append-replace” merges rather than naive append to keep the transcript clean.
- Current observation:
- Hitting Stop can cut off content if the last chunk hasn’t appeared yet; I can lose 30–60+ seconds.
- Expectation (from my perspective):
- Stop should stop capturing more audio but allow all recorded audio to finish transcribing.
- Suggestion:
- Treat Stop as “end of input” while permitting queued processing to complete so nothing recorded is lost.
- I’m open to optional filler-word removal/light smoothing, but would prefer:
- A toggle for aggressiveness.
- Smaller chunks + reconciliation to preserve cadence and detail.
- I’d like to try pro STT models (Perrakey, VoxTroll, Vostrel).
- Opinion:
- Gating OSS models behind Pro (when not hosted by you) can feel off from a user’s perspective.
- Gating cloud inference endpoints makes sense; gating local OSS choice feels less aligned with user expectations.
- The transcript pane becomes a large, fatiguing block of text.
- Suggestions:
- Insert line breaks between sessions/chunks and when I stop/restart.
- Add timestamps per chunk/segment to aid navigation.
- For multi-speaker meetings, longer-term: speaker/time structuring would help.
- I like the audio level animation for feedback.
- The ovular control feels odd to me (subjective).
- No strong alternative proposed—flagging the UX feel.
- It’s unclear whether the autonomy setting applies to transcript, summary, or both.
- Suggestions:
- Independent controls (even if in an “advanced” section), labeled clearly about what each affects.
- I used chat to request “more detail,” which worked, but re-recording/re-generation collapsed the detail back to a high-level structure.
- Suggestions:
- Provide a way to “lock in” desired detail or a persistent instruction so re-gens keep that density.
- Template/system-instruction-level controls for target granularity.
- If I add info later that logically belongs earlier, I’d love the agent to reposition it rather than just append.
- Suggestion:
- Allow non-linear restructuring when the system re-generates the summary so content lands in the right section.
- On the main interface, it isn’t obvious which template is selected for a given note.
- Suggestions:
- Show the selected template prominently (or via hover/label).
- Consider context-aware defaults (e.g., Zoom vs Voice Note may imply different defaults).
- The system-instruction box is ~3 lines tall; modern prompts are often lengthy.
- Suggestion:
- Make the area resizable and/or larger by default.
- Use case: Job Interview template.
- I want to attach a job description and a candidate resume, plus links/docs, without polluting the core note.
- Suggestions:
- Allow attaching files/URLs to a note and/or to a template’s context.
- Clarify where attachments are visible to the model: template-level, note-level, chat-only.
- In chat, support uploads and referencing attachments with scope toggles (template only, this note, summary, etc.).
- Rationale:
- Interviews: role specifics + company criteria.
- Customer work: collateral (e.g., global business plan, MedPIC/Challenger frameworks, contracts/artifacts), so I can analyze across contexts.
- I saw “folders coming”—that would help a lot.
- My workflow:
- I prioritize “Who/What” first, “When” second.
- I often return later to add thoughts to a prior meeting.
- Suggestions:
- Organize by Who/What (company/project/people) then Date.
- Let me add a new session tied to a previous meeting without conflating the original audio timeline.
- Tags across notes (e.g., “Caterpillar” + opportunity number) and the ability to chat across all items with a given tag/folder.
- Defaults: OpenAI, Google Gemini (Flash/FlashLite), OpenRouter.
- Suggestion:
- Consider adding Anthropic as a first-class option (Haiku can be cost-effective here).
- For “Other” endpoints, add guardrails (auto-add v1, handle chat/completions paths) to reduce setup errors for less-technical users.
- Local “Hyper LLM” feels right for responsiveness; perhaps limited for deeper/multimodal tasks.
- Curiosity/questions (not demands):
- Is Hyper Cloud just a bigger/faster variant of the local model, or different models entirely?
- Will there be quotas/usage caps, and multi-model options?
- How do you plan to handle cost variability across users?
- Suggestion:
- A clear reference page/table explaining model choices, strengths, caps, and pricing interactions.
- My high-level summaries may be influenced by my chosen model (Llama 4 Maverick).
- Suggestion:
- Provide defaults that tend to produce balanced detail for note-taking, and allow per-template model/behavior settings.
- Multimodal:
- I’d like to attach slides/images and ask for summaries/analyses.
- Data viz like “chart X from these notes” would be great.
- Voice-to-voice chat:
- With capable STT (e.g., Vostrel Small) and a lightweight TTS (Kokoro, Chatterbox), I could talk to my notes conversationally.
- Longer-term, on-device voice chat feels feasible on modern hardware.
- Chat across folders/tags:
- “Talk to everything tagged Caterpillar,” including attached collateral.
- Obsidian integration fits a privacy-focused audience.
- Additional integrations I’d personally find valuable:
- Notion, Google Docs/Drive, Gmail, OneNote, etc.
- Mobile:
- iOS/Android apps would help for meeting capture and quick notes; desktop-first is okay initially, but cross-device is important.
- “Source analysis” appears to be behind a compute decision; demo/screencap may be outdated.
- If “Ask AI” is coming:
- Align visuals and flows (mental model similar to Google Docs AI helpers).
- Suggestion:
- Keep feature names/status and tutorials in sync to avoid confusion.
- Opinionated take:
- Gating cloud inference behind Pro makes sense.
- Gating locally-run open-source STT models behind Pro feels misaligned (as a user paying to use my own hardware/OSS).
- Pro page copy:
- The “Other small things” bullet undersells value.
- Emphasize: “We include an AI inference runtime (local + cloud options)” if that’s accurate.
- Trials/credits could help (I understand cost variability makes this non-trivial).
- Community access (e.g., Discord):
- I’d keep this open; cultivating champions/templates/integrations/blogs can compound value.
- I’d love docs that mirror Settings and key workflows 1:1.
- Explicit pages for:
- What autonomy affects (transcript vs summary vs both).
- What re-generation does to user-edited summaries.
- The site uses warm peach/cream tones; some screenshots/assets look cold (white/blue frames), reading like dropped-in assets rather than themed.
- Logo kerning/spacing:
- The lightning bolt between R and N is cool; spacing elsewhere feels looser by contrast.
- I’m not a designer; just a subjective note.
- Transcript pane:
- Line breaks between sessions/chunks.
- Timestamps on segments.
- Summary:
- Preserve user-requested detail across re-gens; allow “lock” or persistent directive.
- Template selection:
- Show the selected template on the main list.
- Consider context-aware defaults.
- Attachments:
- Allow attachments (files/links) with scope controls; let chat reference them.
- System instructions:
- Resizable/larger prompt box; support long-form instructions.
- Autonomy:
- Separate controls with clear scope tooltips.
- Stop-recording behavior:
- Finish processing recorded audio before finalizing.
- STT:
- Whisper small is usually fast on Apple Silicon; my hiccups likely relate to execution/timing.
- I’d like to try Perrakey, VoxTroll, Vostrel without Pro gating if they’re local/OSS.
- TTS:
- Pluggable TTS (Kokoro, Chatterbox) could enable voice responses/conversation.
- Local model capability:
- Your local model feels like a sensible speed/size compromise. For bigger tasks (cross-note, multimodal), users may prefer occasional cloud calls.
- Device headroom:
- Even older M1 Pro machines are capable; newer M3/M4 and modern Androids increase headroom.
- Community strategy suggestions:
- Encourage users to build/share templates, write guides, and showcase integrations.
- Highlight community contributions in blogs; consider a templates exchange (not necessarily paid).
- Education:
- Many users are endpoint-/API-setup shy; first-class provider buttons and guardrails help a lot.
- Stopping recording appears to truncate transcript if the last chunk hasn’t emitted yet.
- Suggestion: Let queued processing finish so all recorded audio is transcribed.
- Verbatim transcript with optional light smoothing and controls.
- Smaller chunking (5–10s) with append-replace merging and overlap reconciliation.
- Separate autonomy settings for transcript and summary; clarify scope in the UI.
- Preserve user-directed detail across re-gens; allow “lock”/persistent directives.
- Non-linear summary restructuring (place late info in the right section).
- Template selection visibility; context-aware defaults.
- Resizable system-instructions box; attachments in notes/templates; chat access with scope controls.
- Transcript pane: line breaks, timestamps.
- Add Anthropic as a first-class provider; smooth “Other” endpoint setup.
- Explain Hyper Cloud models, limits, and pricing approach.
- Add mobile apps; broaden integrations (Notion, Google, OneNote, Gmail).
- Multimodal analysis; voice-to-voice chat with TTS.
- Update “Source analysis” visuals; clarify “Ask AI” flows.
- Refine Pro page; reconsider gating OSS STT locally; consider trials.
- Does the autonomy slider affect transcript, summary, or both?
- Can user-set detail preferences persist across re-gens?
- Will Hyper Cloud offer multiple model choices and quotas?
- Roadmap for folders/tags and cross-note chat?
- Plans for multimodal and voice-to-voice as first-class features?
- Address stop-recording truncation so recorded audio always finishes transcription.
- Make the transcript verbatim and adopt smaller-chunk pipeline with clean merging.
- Separate transcript vs summary behavior and autonomy controls; clarify what each affects.
- Preserve user-requested summary detail across re-gens; allow “lock” or persistent directives.
- Show selected template; enlarge/resizable system-instruction editor; enable attachments + scoping.
- Consider ungating local OSS STT models from Pro while keeping Pro gating for hosted/cloud inference.
- Add timestamps/line breaks in the transcript pane.
Overall, I like the product direction and design language, and I appreciate the privacy-first/local inference stance. My biggest personal pain points are transcript fidelity (and not losing content), clarity/control over autonomy and re-generation behavior, and OSS model gating. The areas I’m most excited about are cross-note/tag chat with attachments, multimodal, and voice-to-voice—features that could make Hypernote a central productivity hub.
Please take all of this as one user’s perspective to inform your roadmap—not as a directive.