AI training and your writing

6 min read·Updated 4 months ago

Your writing is yours. Bellwether does not use your manuscripts to train the AI Assistant or any third-party model by default. This article covers exactly what the defaults are, what changes if you opt in, which specific things are never used regardless of opt-in, and how to revoke consent if you change your mind.

The default

AI training is off. This is the default for every account, on every plan.

With the default:

  • Your manuscripts are not used to train any model.
  • Your assistant conversations are not used to train any model.
  • Your cover images are not used to train any model.
  • Your personal preferences and style rules are not used to train any model.
  • Your manuscripts are processed by the AI Assistant when you ask it to, but only for the duration of that request. The processed context is not stored, cached, or reused for any other purpose.

You can verify this at Settings → Privacy. The "Use my writing to improve the assistant" toggle should be off. If it's on, you or someone with access to your account turned it on explicitly.

What's always the case, regardless of opt-in

Some things are never used for training, even if you opt in:

  • Assistant conversations. This is a hard rule. Your chats are retained for your own history (restoring past sessions, auditing) but never fed into any model. Even turning training on doesn't include your conversations.
  • Financial data. Billing history, payment details, invoices, subscription plan.
  • Account metadata. Email, pen names, preferences, style rules, trusted devices.
  • Cover images. Covers are generated by a separate pipeline from the assistant and don't feed back as training data.
  • Comments on retailer-published work. We don't scrape reader reviews to improve the assistant's style guidance.

The only thing that opt-in affects is manuscript content, and only the manuscripts you explicitly mark.

If you opt in

Turning the toggle on doesn't instantly use everything. It shows a per-book checklist:

  • Each book in your account appears.
  • You check the ones you're willing to contribute.
  • Books you don't check are never used.
  • You can change the selection at any time.

Books you opt in to contribute:

  • Contribute their full manuscript text, not just excerpts or summaries.
  • Contribute any metadata attached to the manuscript (genre, description).
  • Contribute to training of the next version of the AI Assistant and related models (blurb generator, metadata rewriter).
  • Are stripped of author-identifying information before training. Your name, email, pen name, and any identifying details that appear inside the manuscript itself are replaced with generic placeholders before the data enters a training run.

Books you don't check are treated exactly as if the toggle were off.

What "contribute" means in practice

Contributed data is used to improve:

  • The AI Assistant's understanding of fiction structure. Pacing, dialogue, character arcs, genre conventions.
  • Style rule enforcement. Learning from patterns across many books what kinds of rules authors commonly set and what those rules look like in practice.
  • Retailer metadata generation. The blurb and description generator improves as the model sees more real published descriptions.
  • Trope detection. Better suggestions for metadata categories and keywords based on book content.

It is not used for:

  • Generating covers (a separate pipeline).
  • Producing direct output of your text in any other user's session. Training a model on your book doesn't make your sentences appear in someone else's draft.
  • Resale or sharing with third parties. We don't sell your data, under any toggle state.

What de-identification means

Before manuscripts enter a training run, we:

  • Replace every instance of your name, email, pen names, and anything in your account profile with generic placeholders.
  • Remove any chapter metadata that could identify you (author notes, dedications referring to specific people).
  • Scan for personally identifiable information (phone numbers, addresses, real names of third parties) and flag for review.

De-identification isn't perfect. A sufficiently determined researcher could in theory re-identify distinctive writing by author. But for practical purposes, your contributed data is anonymised in the training dataset.

If you've written a deeply personal memoir and are worried about even theoretical re-identification, just leave training off for that book. Checked opt-in is book-by-book for exactly this reason.

Revoking

Turn the toggle back off at Settings → Privacy. Or uncheck individual books without turning the whole feature off.

When you revoke:

  • Your data stops being used in future training runs from that moment.
  • Data that was already incorporated into the currently-trained model can't be retroactively removed from that model. Once data is baked into a trained model's weights, it's mathematically entangled with everything else.
  • Expect full propagation through our training pipeline within 72 hours. After 72 hours, any training run that starts does not include your data.

What this means in practice: if a model version was trained with your book's contribution, that version still reflects it. The next model version, trained after your revocation, doesn't include your data. Over time, as models are retrained, your contribution fades.

For accounts that want stronger revocation, Studio plan customers can request a forced retrain notification: when the next training run begins, you're notified so you can verify your data is excluded. Contact support for this.

Assistant conversations and retention

Regardless of training status:

  • Conversations are stored for your history. You can scroll back, resume an old thread, or audit what the assistant has said.
  • You can delete individual conversations under Settings → Privacy → Your AI data → Conversations. Deleted conversations are gone from our systems immediately.
  • You can delete all conversations at once under the same panel. This is immediate and irreversible.
  • Deleted conversations can't be used to train any model from the point of deletion forward (which is moot, since conversations are never used for training anyway, but worth saying).

Conversations are retained indefinitely until you delete them. No automatic expiry.

Third-party processing

The AI Assistant runs on Bellwether's infrastructure with models we license. When you message the assistant:

  • Your message and the relevant manuscript context are processed by our infrastructure.
  • Processing happens in our secured region for your data. US by default, EU available on request for Studio plans.
  • The request is not sent to any third-party service that would use your data for their own training.
  • Our model licenses explicitly prohibit our suppliers from training on customer data.

The specific model names we license and our subprocessors are documented at bellwether.app/legal/subprocessors. Changes to the subprocessor list are announced 30 days in advance via email; you can opt out by cancelling if you disagree with a new subprocessor.

Your data export includes all of this

Everything you've contributed, every conversation, every book is included in the data export (see Exporting your data). If you want a local record of what you've shared before revoking, export before revoking.

The export includes an assistant-history.json with every conversation verbatim and a usage.csv showing every credit spend with a timestamp.

For business customers

If you're on an enterprise agreement (Studio plan with volume discount, or a signed enterprise contract):

  • Your per-organisation policies can override user-level defaults.
  • Your admin can disable training globally across all users in the organisation.
  • Your admin can require EU-region processing.
  • Your admin can require a specific data processing addendum.

Individual users can't change organisation-level policies. Contact your admin if organisation-level settings are blocking something you want to do.

Common edge cases

I turned the toggle on by mistake. Turn it back off. Anything not actively in a training run (99%+ of the time) stops being used immediately. Already-trained models cannot be modified, but your data stops contributing to future ones.

I want to confirm specific books are not used. Open Settings → Privacy → Per-book consent. Each book shows its current status. Unchecked books are never used regardless of the main toggle.

I signed up for an enterprise account and training is on by default. Company-level agreements sometimes reverse defaults. Check with your admin. You can't change organisation-level policies from the user settings, but your specific opt-outs are respected within the policy.

I deleted my account. Are previous contributions erased? Your data is deleted from our systems after the 30-day grace period. Contributions to already-trained models cannot be retroactively removed, but no future training run after deletion includes your data.

My writing is similar to another author's whose data is in the training set. Possible, and the model won't plagiarise them into your draft. Trained models don't reproduce specific training examples verbatim; they learn patterns at a level of abstraction well above exact sentences.

Can I request a forced retrain today to exclude my data? Forced retrains happen on our schedule (every few months per model). You can request notification of the next one. Ad-hoc retrains on user request aren't offered; the energy and cost would make them impractical.

What about model fine-tuning on my own data for personalised output? A separate feature, on the roadmap for Studio plans. Per-user fine-tuning with opt-in data would give you an assistant trained specifically on your voice. Not available yet.

Was this helpful?

Need more help? Contact support.