TSV to Markdown: Convert Tab Data & Paste Online (or Code)
Quick answer: TSV to Markdown conversion turns tab-delimited text (\t) into a GitHub Flavored Markdown (GFM) pipe table: one row per line, cells separated by |, and a divider row under the header. For a one-off table, paste your cells into the free TSV to Markdown converter and copy the result. No install, no upload.
Raw TSV (tab characters between columns, commas inside cells, leading zeros that must survive):
sku reagent supplier_note unit_price
0042 Taq polymerase, 500 U Smith, J.; Lot 7731 129.00
0137 dNTP mix (10 mM) Ships on dry ice 48.50
0201 Primer set 16S/ITS Müller & Söhne, GmbH 212.75
GFM pipe table output:
| sku | reagent | supplier_note | unit_price |
| ---- | --------------------- | -------------------- | ---------: |
| 0042 | Taq polymerase, 500 U | Smith, J.; Lot 7731 | 129.00 |
| 0137 | dNTP mix (10 mM) | Ships on dry ice | 48.50 |
| 0201 | Primer set 16S/ITS | Müller & Söhne, GmbH | 212.75 |
Column padding is cosmetic. GFM only needs the pipes and the divider row.
Which method fits your situation
| Method | Speed | Dependencies | Clipboard-ready | Best for |
|---|---|---|---|---|
| Online tool (md-convert.org) | Instant | None (browser) | Yes, paste directly | One-off tables, READMEs, PR descriptions |
| Python + Pandas | Fast once scripted | Python, pandas, tabulate | Indirect (read_clipboard) |
Data pipelines, notebooks |
| Pandoc CLI | Fast | Pandoc 3.x | No, file in, file out | CI/CD doc builds |
| AWK | Instant startup | None on Linux/macOS | Via pbcopy / xclip pipe |
Quick terminal jobs, simple data |
The Hidden TSV Reality: Why Pasted Spreadsheet Cells Break in Markdown
When you select cells in Excel, Google Sheets, or Numbers and press Ctrl+C, the app writes several clipboard formats at once: HTML, rich text, and plain text. The plain-text flavor is TSV. Cells are separated by a tab character, rows by a newline. That is why "copy Excel, paste Markdown" gives you a block of text with invisible gaps instead of a table. Markdown has no idea that those gaps are column boundaries, and a GFM table needs pipes plus a divider row before any renderer treats it as a table.
Spreadsheets also apply CSV-style quoting to tricky cells. If a cell contains a tab, a line break, or a double quote, the app wraps it in quotes and doubles any inner quotes. Splitting on \t alone will mangle those cells, so any converter you trust has to parse quoting rather than split blindly.
The Git and documentation advantage. An .xlsx attachment in a pull request is a binary blob. Nobody can review it inline, and git diff shows nothing useful. A pipe table is plain text. If someone changes the price of SKU 0137 from 48.50 to 52.00, the diff is one line, reviewers can comment on that row, and git blame tells you who changed it. Tables stay close to the prose they explain, in READMEs, wikis, docs sites, and static site generators.
Why bioinformatics lives on TSV. BED, GFF, VCF, sample sheets, and count matrices are tab-delimited. Annotation fields routinely contain commas (hemoglobin subunit beta, adult), and tabs almost never appear in free text. CSV forces quoting around those fields. TSV usually doesn't, so a file's structure survives hand-editing and shell tools. The cost shows up when you want that data in documentation, because Markdown has no native import. If your source uses commas, the CSV to Markdown guide covers the quoting rules for that format.
Method 1: Instant Browser Conversion with md-convert.org
For a single table, the fastest route is to convert TSV to Markdown table online without leaving the browser.
- Copy the cell range in your spreadsheet, or have a
.tsvfile ready. - Paste into the input area, or upload the file.
- Set your options: toggle the header row, choose column alignment (
:---left,:---:center,---:right), and enable pipe escaping so a literal|in a cell becomes\|instead of splitting the column. - Copy the Markdown output into your document.
Privacy. Conversion runs as JavaScript in your browser. Pasted data and uploaded files are never sent to a server, and the site records only cookieless, aggregate usage metrics. Don't take that on trust: open DevTools (F12), switch to the Network tab, run a conversion, and check that no request carries your data. That matters when the table holds unpublished sample IDs, customer names, or internal pricing.
The converter is one of 16 on the all-in-one Markdown conversion workspace (16 dedicated converters). If your data starts life in a different format, the CSV to Markdown converter and the Excel to Markdown converter handle those inputs directly.
Method 2: Convert TSV to Markdown Programmatically (Python & Pandoc)
Scripts make sense when the conversion is repeatable: a nightly report, a docs build, a notebook.
Python with Pandas and Tabulate
pip install pandas tabulate
import pandas as pd
df = pd.read_csv(
"data/samples.tsv",
sep="\t",
dtype=str,
keep_default_na=False,
)
# to_markdown() does not escape pipes or newlines, so do it first
df = df.apply(
lambda col: col.str.replace("|", r"\|", regex=False)
.str.replace("\n", "<br>", regex=False)
)
print(df.to_markdown(index=False, disable_numparse=True))
Three arguments do the real work here:
dtype=strstops type inference. Without it, pandas reads SKU0042as the integer42and a sample ID like5E10as the float50000000000.0. Once the zeros are gone, you can't recover them.keep_default_na=Falsekeeps literal strings such asNA(Namibia's country code, or "not available" in a lab sheet) from silently turning into empty cells.disable_numparse=Truetells Tabulate not to reformat numeric-looking strings. Without it, a price of19.90can come out as19.9.
Pure Python (Standard Library Only)
No dependencies, which suits locked-down servers and CI images. Requires Python 3.9+.
#!/usr/bin/env python3
"""tsv2md.py: TSV to GFM pipe table, standard library only."""
import csv
import sys
def clean(cell: str) -> str:
cell = cell.replace("|", r"\|")
return cell.replace("\r\n", "<br>").replace("\n", "<br>").strip()
def tsv_to_markdown(path: str, encoding: str = "utf-8-sig") -> str:
with open(path, newline="", encoding=encoding) as f:
rows = [[clean(c) for c in row] for row in csv.reader(f, delimiter="\t")]
rows = [r for r in rows if any(r)]
if not rows:
return ""
width = max(len(r) for r in rows)
rows = [r + [""] * (width - len(r)) for r in rows]
header, *body = rows
lines = [
"| " + " | ".join(header) + " |",
"| " + " | ".join(["---"] * width) + " |",
]
lines += ["| " + " | ".join(r) + " |" for r in body]
return "\n".join(lines)
if __name__ == "__main__":
if len(sys.argv) != 2:
sys.exit("usage: tsv2md.py input.tsv")
print(tsv_to_markdown(sys.argv[1]))
csv.reader with delimiter="\t" handles quoted cells and embedded line breaks, which str.split("\t") cannot. The utf-8-sig encoding strips a byte-order mark if Excel added one. Every value stays a string, so 0042 stays 0042.
The Pandoc Power Command
pandoc -f tsv -t gfm input.tsv -o output.md
Pandoc is the right choice when Markdown is a build artifact, not something you edit by hand. A Makefile rule regenerates the table whenever the data changes:
docs/samples.md: data/samples.tsv
pandoc -f tsv -t gfm $< -o $@
The TSV reader only exists in recent Pandoc 3.x releases, so run pandoc --list-input-formats | grep tsv before wiring it into a pipeline. Pandoc also expects UTF-8 and treats " as a quote character, which matters for the hazards covered below.
Method 3: Command Line & IDE Automation (Terminal & VS Code)
AWK one-liner. AWK ships with Linux and macOS, and -F'\t' sets the tab delimiter:
awk -F'\t' '
{ for (i = 1; i <= NF; i++) gsub(/\|/, "\\|", $i)
printf "|"; for (i = 1; i <= NF; i++) printf " %s |", $i; printf "\n"
if (NR == 1) { printf "|"; for (i = 1; i <= NF; i++) printf " --- |"; printf "\n" } }
' samples.tsv | pbcopy
Swap pbcopy for xclip -selection clipboard on Linux or clip on Windows. The pipe-escaping gsub is the part most copy-pasted snippets leave out. AWK still breaks in two places: quoted fields containing line breaks (it reads one line at a time, so a multiline cell becomes two broken rows) and quoted fields containing tabs. For data like that, use the Python script.
VS Code workflow. Paste the TSV into a scratch file, then use regex Find and Replace (Alt+R toggles regex):
- Replace
\twith|. - Replace
^(.*)$with| $1 |. - Insert a divider row under the header, such as
| --- | --- | --- | --- |.
If you convert often, marketplace extensions such as Excel to Markdown table automate the paste step, and Markdown All in One realigns an existing table on save. Both help with formatting but won't fix pipes or multiline cells in your source data.
How to Import and Render TSV in R Markdown
R Markdown renders tables from data frames, so the conversion happens inside the document:
```{r samples-table, echo=FALSE, message=FALSE}
library(readr)
df <- read_tsv(
"data/samples.tsv",
col_types = cols(.default = col_character())
)
knitr::kable(df, format = "pipe")
The base R equivalent is `read.delim("data/samples.tsv", stringsAsFactors = FALSE, colClasses = "character")`. As in pandas, forcing character columns preserves leading zeros.
Two practical notes. First, knitr evaluates chunks with the `.Rmd` file's folder as the working directory, so `data/samples.tsv` must be relative to the document, not to your RStudio project root or console directory. Second, `echo=FALSE` hides the code and shows only the table, which keeps client-facing reports clean. `format = "pipe"` makes `kable` emit the same GFM syntax shown earlier.
## Markdown to TSV: Extracting Tables Back to Spreadsheets
Sometimes the data lives in a README and someone wants it back in a spreadsheet. If you've searched for "markdown to tsv" or "markdown table to tsv", this script covers it. It reads the first GFM pipe table in a file, skips the alignment row (`|---|:---:|`), splits only on unescaped pipes, turns `\|` back into `|`, and converts `<br>` back into line breaks.
```python
#!/usr/bin/env python3
"""md2tsv.py: first GFM pipe table in a file to TSV, standard library only."""
import argparse
import csv
import re
import sys
SEPARATOR = re.compile(r"^\s*\|?\s*:?-+:?\s*(\|\s*:?-+:?\s*)*\|?\s*$")
UNESCAPED_PIPE = re.compile(r"(?<!\\)\|")
LINE_BREAK = re.compile(r"<br\s*/?>", re.IGNORECASE)
def split_row(line: str) -> list[str]:
line = line.strip()
if line.startswith("|"):
line = line[1:]
if line.endswith("|") and not line.endswith("\\|"):
line = line[:-1]
return [
LINE_BREAK.sub("\n", cell.strip().replace("\\|", "|"))
for cell in UNESCAPED_PIPE.split(line)
]
def extract_table(text: str) -> list[list[str]]:
rows, n = [], 0
for line in text.splitlines():
if not line.strip().startswith("|"):
if n:
break # first table has ended
continue
n += 1
if n == 2 and SEPARATOR.match(line):
continue # alignment row
rows.append(split_row(line))
return rows
def main() -> None:
parser = argparse.ArgumentParser(description="Markdown table to TSV")
parser.add_argument("input")
parser.add_argument("-o", "--output", help="write to a file instead of stdout")
args = parser.parse_args()
with open(args.input, encoding="utf-8") as f:
rows = extract_table(f.read())
if not rows:
sys.exit("No pipe table found.")
if args.output:
with open(args.output, "w", newline="", encoding="utf-8") as out:
csv.writer(out, delimiter="\t", lineterminator="\n").writerows(rows)
else:
sys.stdout.reconfigure(encoding="utf-8")
csv.writer(sys.stdout, delimiter="\t", lineterminator="\n").writerows(rows)
if __name__ == "__main__":
main()
Run it and pipe straight to the clipboard, then press Ctrl+V in Excel or Google Sheets:
python md2tsv.py README.md | pbcopy # macOS
python md2tsv.py README.md | xclip -selection clipboard # Linux
python md2tsv.py README.md -o table.tsv # any OS, then open the file
csv.writer quotes only the cells that need it (tabs, line breaks, quotes), which is the same convention spreadsheets use when reading pasted text, so multiline cells land in a single cell.
One trap on the way back: Excel and Google Sheets guess types on paste, so 0042 becomes 42. Before pasting, format the destination columns as text (Excel: Format Cells → Text; Google Sheets: Format → Number → Plain text).
The TSV Edge-Case Test Matrix
TSV removes the comma problem, but five other hazards still break conversions. I tested each against the five common approaches. ✅ means handled by default or with the script above, ⚠️ means it needs a workaround, ❌ means it breaks.
| Hazard | md-convert.org | Pandas | Python stdlib | Pandoc | AWK |
|---|---|---|---|---|---|
| Embedded commas in text | ✅ | ✅ | ✅ | ✅ | ✅ |
| Pipe ` | ` inside cells | ✅ (pipe escaping option) | ⚠️ escape first | ✅ | ✅ |
| Multiline cells | ⚠️ check the preview | ⚠️ replace \n with <br> |
✅ (<br>) |
⚠️ verify output | ❌ |
Leading zeros (007, 0142) |
✅ | ⚠️ dtype=str |
✅ | ✅ | ✅ |
| UTF-16 LE input | ⚠️ re-save as UTF-8 | ✅ encoding="utf-16" |
✅ encoding="utf-16" |
❌ iconv first |
❌ iconv first |
Workarounds:
- Embedded commas: nothing to fix. This is why TSV beats CSV for descriptive text.
- Pipes in cells: escape as
\|before output. Pandas needs thestr.replaceshown earlier, and AWK needs thegsub. - Multiline cells: a GFM table row must stay on one line, so replace internal newlines with
<br>. GitHub, GitLab, and most static site generators render it as a line break. If you use AWK and have multiline cells, switch to the Python script. - Leading zeros: keep every column as a string at read time. In pandas that's
dtype=str; in R it'scolClasses = "character"orcol_character(). - UTF-16 LE: Excel's "Unicode Text" export is UTF-16 LE with a byte-order mark. Convert before using UTF-8-only tools:
iconv -f UTF-16 -t UTF-8 export.txt > export.tsv. If Japanese text turns into garbage, the file is probably CP932 (Shift_JIS) from a Japanese-locale Windows machine, so useiconv -f CP932 -t UTF-8instead.
One more case, common in bioinformatics: a stray double quote such as 5" tubing can make quote-aware parsers (Pandas, Pandoc, csv.reader) swallow the following rows into one cell. For files that never use quoting, pass quoting=csv.QUOTE_NONE in Python or quote="" in R's read.delim.
Frequently Asked Questions
What is the fastest way to convert copied Excel cells into a Markdown table?
Copy the cell range in Excel, Google Sheets, or Numbers, then paste it into a browser-based converter such as md-convert.org. The clipboard already holds tab-separated text, so the tool turns it into a pipe table immediately. Copy the output into your README or docs. No install and no upload needed.
Can I convert a Markdown table back into TSV for Google Sheets?
Yes. Run the Python script in the Markdown to TSV section: it skips the alignment row, unescapes \|, and writes tab-separated output. Pipe the result to your clipboard (pbcopy, xclip, or clip) and press Ctrl+V in a sheet. Format target columns as plain text first to keep leading zeros.
How do I handle TSV files with Japanese or non-ASCII characters?
Save the file as UTF-8 and everything passes through untouched. Excel on Japanese Windows often exports CP932 (Shift_JIS), and its "Unicode Text" option writes UTF-16. If you see garbled characters, convert first with iconv -f CP932 -t UTF-8 in.txt > out.tsv, or use -f UTF-16 for Unicode Text exports.
Does md-convert.org store or read my uploaded TSV data?
No. The converter runs as JavaScript in your browser, so pasted text and uploaded .tsv files never leave your device. You can verify it: open DevTools, switch to the Network tab, run a conversion, and confirm no request carries your data. The site records only cookieless, aggregate usage metrics.
Next step: for a table you need right now, paste it into the free TSV to Markdown converter. For anything repeatable, use the Pandas, stdlib, or Pandoc patterns above, and keep md2tsv.py handy for the return trip.
