cat and tee

tee redirects stdin to stdout and files specified as arguments.

    exec py -c print("hello") | 
      tee log1 log2
  

To also capture stderr, redirect stderr to stdout prior to piping tee.

    exec 2 2>&1 | tee path/to/logfile
  

sort and uniq

sort and uniq -c are combined to perform a "groupby-count" operation on lines of stdin. In the following snippet, the count of html pages per parent directory is logged to stdout. The stage of sort is needed to prepare groups, since uniq -c counts non-unique contiguous lines.

  find -type f |    # find files
  xargs dirname |   # select parent
  sort | uniq -c    # groupby occurence
  

It is possible to sort -u, which is equivalent to sort | uniq, on most POSIX systems.

uniq -c showing counts first, it can be ordered by a follow-up sort -hr. h sorts by human numeric count. Without it, 11 goes before 2, 131 before 14, etc. Showing outliers can be done by limiting with head -n for the top n most common, tail -n for the top n rarest.

  find -type f \
  -name '*.html' |        # find html
  xargs dirname |         # select parent
  sort | uniq -c |        # groupby occurence
  sort -hr |              # orderby most
  head -n15               # show top 15
  

grep

Its name comes from the ed command g/re/p (global regular expression search and print), which has the same effect. Examines each line of data it receives from standard input and outputs every line that contains a specified pattern of characters.

fmt and pr

fmt formats stdin to stdout. pr paginates stdin inserting footers, page breaks, headers, for printing out.

head and tail

Head and tail output resp. the beginning and ending of stdin to stdout, cutting the rest. Python piped in head/tail throws if not explicitely handling the SIGPIPE signal.

    from signal import signal, SIGPIPE, SIG_DFL
    signal(SIGPIPE, SIG_DFL)
    

tail -n+M removes the first M lines and head -n-M removes the last M lines.

    
    # Example: extracting the core of tree -H .
    # 1. Removes the first 28 lines
    # 2. Removes the last 9 lines
    tree -H . | tail -n+28 | head -n-9
  

tr and sed

Changing patterns from stdin lines into stdout. Useful for casing, field separation, line separation changes, tr has a simple syntax, while sed leverages regex to perform any stream editing. "html cannot be parsed with regex", but when processing your own mark-up, most jobs can be done with sed the quickest. To perform line-wise prepending, appending, or "circumpending", sed is one of the fastest ways, matching regex ^ and $. Prepending to all stdin lines: sed -e 's!^!&foo!g'
Appending to all stdin lines, before new-line/carriage-return: sed -e 's!$!&bar!g'. Escaping parenthesis like \(pattern\) enables matching a group, then usable in replacement via \1 for group 1 and so on. For example, sed 's/\(.*\)/[\1]/' wraps lines with brackets. The flag -z on GNU sed enables replacing newlines, matching on \n. By default not possible since sed works line-by-line. The amperstand & un-escaped enables concatenations instead of replacements. For example, sed 's!class="foo!& bar!g' will add bar after each occurence of class="foo.

sed ... p: prints.

sed ... d: deletes. Can be seen as the "inverting" flag, close in spirit to the grep -v

sed ... g: replaces all vs. replacing one.

      echo 0123 | sed 's [0-9] N ' # N123
      echo 0123 | sed 's [0-9] N g' # NNNN
    

Ways to convert stdin paths into anchors

  # via globbing, awk
  ls path/to/dir | 
  awk '{printf "<a href=\"%s\"></a>\n", $0}'
  # via ls, xargs, printf
  ls path/to/dir | 
  xargs -I{} printf '<a href="%s"></a>\n' {}
  # via echo, printf
  echo path/to/dir/* | 
  xargs printf '<a href="%s"></a>\n'
  # via sed - not great due to the forward flash
  ls path/to/dir | 
  sed 's|.*|<a href="&"></a>|'
  # as an alias, escaping quotes
  alias anchors=\
  'sed '\''s|.*|<a href="&"></a>|'\'''